Claude Transcripts docs GitHub
Work in progressUnder active development — not tested as ready for use. Breaking changes land without notice, stored data may need to be discarded between revisions, and there is no auth or security model.

27. Full-content chunks in CouchDB (per-turn content, not just byte ranges)

Date: 2026-07-22

Status

Accepted, implemented on the write and read paths. The hook and backfill write entries[] by default, validated at the webapi when ingested. Consumers: speaker_split/by_role (v4, GET /api/sessions/{id}/turns), speaker_split/by_role_time (v5, GET /api/turns), chunks/entries_by_session (v6, which makes chunks the default source for GET /api/sessions/{id}/transcript and narrows ADR 0014), and Meilisearch's turns index (ADR 0009).

Sessions adopted before this, or with --no-content, have byte-range-only chunks; they fall back to S3 on read and are absent from content search. backfill --force rebuilds them (--replace-live for live-recorded ones). The migration proposed below to re-derive chunks from S3 was not built.

Context

A session is stored two ways. The byte-faithful transcript lives in S3 as a single transcript.jsonl object (ADR 0014), and CouchDB holds append-only metadata: event markers, a summary doc at SessionEnd, and — added since — chunk docs written mid-flight for crash resilience (mid-flight-chunking.md).

The mid-flight chunking mechanism is already built: the hook flushes chunk docs as the session grows, ingest is idempotent with stable ids (chunk:<session>:<byteStart>), and backfill reconstructs the same chunks for adopted history. But a chunk doc today records only a byte-range slice — byte_start, byte_end, entry_count — a pointer into the S3 transcript. It does not contain the messages.

The consequence: CouchDB cannot see an individual turn. Anything that needs per-turn structure — speaker-split views (user vs Claude), per-turn search indexing, map-reduce feature extraction, prompt/instruction provenance — is blocked, because the only place the turns exist in parsed form is the S3 blob, and map-reduce can't run over S3. The couchFullContentChunks feature flag was added in anticipation of this but nothing populates content yet.

This is the remaining half of the logging rework (roadmap #4).

Decision

Promote chunk docs from byte-range pointers to full-content chunks: parse the transcript entries and store their content in CouchDB, so map-reduce views operate directly on turns.

Consequences