> ## Documentation Index
> Fetch the complete documentation index at: https://surf-dcinside-api.kr.ask.surf/llms.txt
> Use this file to discover all available pages before exploring further.

# Event Deduplication

> Proposed LLM same-event rankings with explicit published-feed context and incomplete coverage

The API performs LLM same-event deduplication after raw scoring and before committing the cycle response. Reports of the same event from different galleries count as duplicates. Posts about the same person, subject or category can still describe different events. This extension proposes a Surf ranking aid; the client retains the hard publishing gate.

The provider defaults to `gpt-5.6-sol` through Surf's pure Codex OAuth Responses route. There is no enterprise API-key or hybrid fallback. This provider migration does not enable the deployment toggle or change the event/grouping rules, raw scores, or client publishing gate.

## Published feed context

Supply `published_windows.main` and `.light` independently. Each array is the complete rolling window of at most 20 actually published posts for that tier, current as of `cycle_time`. The API evaluates the same pending candidate pool against both windows; candidates do not need a new target-tier field.

```json theme={null}
{
  "published_windows": {
    "main": [{
      "gall_id": "news_gallery",
      "post_no": 120,
      "placed_at": "2026-09-30T15:20:00+09:00"
    }],
    "light": []
  }
}
```

Only `gall_id`, `post_no` and `placed_at` are required per item. `placed_at` requires an explicit offset and cannot be after `cycle_time`. Repeated post identities within one tier or more than 20 entries are rejected. The same post may appear in both tiers, and a candidate already in a supplied feed can be marked as a published duplicate.

Candidate text is sent once. The API archives its received title/body/image metadata durably and retains it after a saved, hidden, deleted or aged-out removal. An explicit published reference reuses that content and existing image uploads, so the snapshot does not resend text or JPEG bytes. Archiving or removing a post never makes it a publication.

For a post whose content was never delivered, optionally supply `title`, `body_html`, `images_in_post` and `image_positions` once as bootstrap evidence; later snapshots can use references alone. Bootstrap fills missing archive fields and cannot replace received candidate evidence. Existing pending records are backfilled when upgrading; already-removed legacy records without retained content cannot be reconstructed. Unknown references or missing evidence make the relevant check incomplete/unchecked rather than inventing content. The archive has no new age expiry or candidate quota.

Omitting a tier or using `null` means its feed context is unknown. An empty array means the feed is known to be empty. Missing context is labelled `candidate_only` and cannot produce a completed feed check. The server does not reconstruct actual publication from `saved` removals, recommendations or outcomes. It does not append recommendations to either window or reuse a previous cycle's snapshot when the current one is absent.

The sender is responsible for snapshot completeness and freshness as of the cycle. An old `placed_at` does not by itself mean the snapshot is stale: a quiet feed can contain old posts. There is no invented age expiry. If the client cannot supply a current complete snapshot, omit that tier and retain the client-side rolling gate. This API cannot independently verify that a client-provided window contains every actual publication.

## Read the response

`results` retains every original raw score row and its provenance. Deduplication does not change a score, drop a row, impose a per-cycle selection quota, or publish anything. `deduplication.tiers.main` and `.light` each contain:

| Field | Meaning |
| - | - |
| `context` | `published_window` when a snapshot was supplied; `candidate_only` otherwise. |
| `ranked` | Checked eligible representatives ordered by descending raw score, then gallery and post number for ties. |
| `duplicates` | Excluded post references, their representative/published reference in `duplicate_of`, and `same_event_candidate` or `same_event_published`. |
| `unchecked` | Candidates without a completed semantic assessment. They are not certified unique. |
| `warnings` | Reasons for missing context, content evidence, provider errors or incomplete work. |

Every scored candidate appears exactly once in that tier's ranked, duplicates or unchecked lists. Non-scored rows remain only in the original results. A duplicate of a published Main post is excluded for Main without automatically being excluded for Light.

| Status | Meaning |
| - | - |
| `completed` | The supplied tier snapshot and all scored candidates were checked. This is a model assessment, not a guarantee of semantic accuracy. |
| `partial` | Some context or comparisons are missing. Any returned ranking has only the reported coverage; candidate-only rankings do not satisfy the published last-20 gate. |
| `unavailable` | No dedup ranking can be certified, including provider/configuration/deadline failures. Raw scores remain available. |
| `disabled` | Deduplication was explicitly disabled in deployment configuration. |

The report records model/prompt provenance and coverage counters. Check both each tier's status and its context; the aggregate status does not erase tier differences. Unavailable/disabled tiers have empty ranked and duplicate lists, with all scored posts unchecked.

## Missing images, caching and time limits

Images may arrive one or two cycles after candidate text. Raw scoring never waits for them. Deduplication reads only uploaded immutable local bytes and never fetches URLs from HTML. Missing or oversized evidence and incomplete comparisons are reported; image-only posts without usable evidence stay unchecked. Text-bearing posts can be assessed only to the extent supported by the evidence and report.

Validated LLM decisions are cached by content/image fingerprints and model/prompt configuration. Late image uploads change those fingerprints for future cycles. Requests are split by engineering payload budgets, without a business limit on the number of candidates; comparisons unfinished within the deadline remain unchecked.

The client allows 60 seconds from sending including at most one retry. `RTB_RESPONSE_BUDGET_SECONDS` defaults to 25 seconds from server request-handler entry; after raw scoring, only the remainder is available to LLM deduplication. This controls dedup work and does not preempt the existing scorer or prove the external deadline. Full-queue latency and coverage must be verified in the deployed environment.

The final report is persisted with its raw scores. Same-ID POST retries and GET return the original completed response without another LLM call, even after provider recovery or later images. A new cycle can produce a new report. Older completed cycles may omit `deduplication` entirely.
