Hierarchical Document Splitting¶
Split content at Markdown headings, retaining section structure and source information.
Use it for text with heading levels. To control chunk length in characters, use Segment.
Prepare the input¶
Bind upstream documents. Headings must appear as Markdown heading lines, such as lines starting with # or ##. Visually bold text alone does not create a heading boundary.
Headings determine split positions. There are no chunk-length or overlap settings.
Splitting behavior¶
The node recognizes Markdown code fences and does not treat # inside them as headings. Code-type records remain intact.
Long sections can therefore produce long chunks. Add length-based chunking afterward if needed, and check that the required structure is preserved.
Use the output¶
The main output is documents, retaining source metadata and adding position fields such as parent_index, chunk_index, and chunk_type.
These fields relate chunks to parent documents. Continue binding documents downstream.
Example: index manual sections¶
Flow: General Document Parsing → Hierarchical Document Splitting → Embed and Add to Knowledge Base.
Check that parsed text contains Markdown headings before splitting. For unusually long sections, add length-based chunking and inspect a sample.