AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→AI Engineering Program — go from software engineer to production AI engineer · Live training with Kirill Eremenko · Watch the program breakdown→

Q: What is chunking, and why does bad chunking ruin RAG?

Chunking means splitting your documents into pieces before you embed them. Those pieces (chunks) are what gets stored, searched, and handed to the model as context.

Why split at all? Because each chunk gets one embedding, one vector of meaning. A whole document contains dozens of ideas, so one vector for all of it is a blur; retrieval works on focused chunks about one thing.

Now the failure. Chunk carelessly (say, cutting every 500 characters) and you slice through the middle of sentences, lists, and ideas. A chunk that ends mid-thought has a distorted meaning, so its embedding is distorted too. Then retrieval fetches fragments, the model answers from incomplete context, and quality drops silently. That's the frustrating signature of bad chunking: the pipeline runs perfectly, the right document is even found, and the answers are still wrong or shallow.

The fix is chunking that respects structure: split on paragraph and section boundaries, keep lists and tables intact, and add overlap between chunks so ideas that straddle a boundary appear whole in at least one chunk.

← Back to the full FAQ