How you cut a document sets the ceiling on every answer drawn from it. It is the first production bottleneck and the largest quality lever — and it is usually an afterthought.

Splitting a document into pieces looks like plumbing, and it is treated accordingly: pick a chunk size, add some overlap, move on to the interesting parts. That ordering is backwards. Whatever the cut destroys is unavailable to every later stage — no reranker, no larger model and no better prompt recovers a distinction that was severed before retrieval began.
The two failure modes are symmetrical and both are quiet. Cut too small and each piece loses the context that made it meaningful; the retriever returns fragments that are individually on-topic and collectively insufficient. Cut too large and precision collapses: the relevant sentence arrives buried in padding, competing for attention with everything else in the chunk, and costing tokens for the privilege. Neither produces an error message. Both produce answers that are merely worse.
The deeper point is that a document usually already has structure — headings, sections, tables, a table of contents someone thought about. Fixed-size character splitting deliberately discards that, then asks similarity search to reconstruct it statistically. Splitting along the existing structure instead, so a section stays a section and a table stays a table, keeps the author's own organisation intact and does the retriever's work for it. For long, well-structured documents this is the difference between navigating and guessing.
Which makes chunking a place where the general thesis shows up concretely: structure that already exists is worth more than structure inferred later. A markdown file with headings is not raw text awaiting processing — it is a document that has already been cut, by a human, along meaningful lines. The engineering task is to preserve that, not to overwrite it with an arbitrary character count.