Skip to content

fix(md): keep every line of pipe-led text that has no delimiter row - #4454

Open
breken-ai wants to merge 1 commit into
docling-project:mainfrom
breken-ai:fix/md-table-without-delimiter-row
Open

breken-ai wants to merge 1 commit into
docling-project:mainfrom
breken-ai:fix/md-table-without-delimiter-row

Conversation

@breken-ai

Copy link
Copy Markdown

The Markdown backend drops a line of text when a paragraph starts with a pipe but has no GFM delimiter row.

A line starting with | opens a table in _iterate_elements. When _close_table builds it, it always throws away the second buffered row, assuming it is the |---|---| delimiter row. If there is no delimiter row, that second row is real content and it disappears. The remaining lines turn into a table that was never in the source.

This shows up with Pandoc line blocks, which are used for addresses and verse, and with pipe tables written without a delimiter row:

Ship to:

| Jane Doe
| 221B Baker Street
| London NW1 6XE

On main, export_to_markdown() gives:

Ship to:

| Jane Doe       |
|----------------|
| London NW1 6XE |

The street line is gone. | a | b | / | 1 | 2 | / | 3 | 4 | loses 1 | 2 in the same way.

Change: GFM only recognizes a table when a delimiter row follows the header. The pipeless path (_starts_pipeless_table) already requires one. _close_table now applies the same rule: if the second buffered row is not a delimiter row, the lines are added as the paragraph they are, so every line is kept. Real tables go through the existing code unchanged.

With the fix the example above becomes | Jane Doe | 221B Baker Street | London NW1 6XE as one text item.

Tests: test_convert_leading_pipe_lines_without_delimiter_row_keep_all_text covers a delimiter-less pipe table and a pipe-led line followed by plain text. It fails on main (a table is built and a line is lost) and passes with the fix. tests/test_backend_markdown.py, tests/test_backend_html.py, tests/test_backend_csv.py and tests/test_backend_asciidoc.py pass (113 passed, 1 skipped) and no reference data changed. ruff check, ruff format --check and ty check on the touched files are clean. ty shows the same two warnings as on main.

I found and fixed this with help from an AI coding assistant, and I reviewed the change and the test results myself.

Checklist:

  • Documentation has been updated, if necessary.
  • Examples have been added, if necessary.
  • Tests have been added, if necessary.

The Markdown backend opens a table as soon as a line starts with a pipe
and, when it closes the table, always discards the second buffered row
as the GFM delimiter row. When there is no delimiter row, as in a Pandoc
line block or a pipe table written without one, that second line is real
content and it was silently dropped, while the rest became a one-column
table.

GFM only recognizes a table when a delimiter row follows the header, and
the pipeless path already requires one. Apply the same rule here: without
a delimiter row, emit the buffered lines as the paragraph they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: breken-ai <312387581+breken-ai@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @breken-ai, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 83.33333% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
docling/backend/md_backend.py 83.33% 0 Missing and 1 partial ⚠️

📢 Thoughts on this report? Let us know!

@ceberam ceberam left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @breken-ai for spotting this issue. The bug is real and the diagnosis is accurate.
The fix is correct, well-scoped, and does not regress existing behavior.
Please, address the inline comment and the PR will be ready to be merged.

Comment on lines +350 to +352
# GFM: a table needs a delimiter row right after its header, so
# lines that merely start with a pipe are a paragraph. Building a
# table here would drop the second line, taken for that row.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The first sentence is useful: it states the GFM rule. The second sentence explains what the old, buggy code did, not what a future maintainer needs to know. It should be removed or rewritten to describe the invariant. For example:

Suggested change
# GFM: a table needs a delimiter row right after its header, so
# lines that merely start with a pipe are a paragraph. Building a
# table here would drop the second line, taken for that row.
# GFM: a table requires a delimiter row immediately after its header.
# Without one, lines that start with a pipe are ordinary paragraph text.

@ceberam

ceberam commented Oct 1, 2026

Copy link
Copy Markdown
Member

@breken-ai Please check the PR #4488
I have refactored the table parsing to leverage marko GFM extension. Instead of a manual string manipulation to parse tables (and the challenges to deal with complex edge cases like this one), we can leverage the Markdown(extensions=["gfm"]):

  • marko natively understands the GFM Table specification.
  • The parser itself produces a clean, structured tree of table nodes.

I checked and with that approach the case in this PR is no longer an issue, since marko does not recognize does Pandoc line blocks as tables.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants