Conversation
The Markdown backend opens a table as soon as a line starts with a pipe and, when it closes the table, always discards the second buffered row as the GFM delimiter row. When there is no delimiter row, as in a Pandoc line block or a pipe table written without one, that second line is real content and it was silently dropped, while the rest became a one-column table. GFM only recognizes a table when a delimiter row follows the header, and the pipeless path already requires one. Apply the same rule here: without a delimiter row, emit the buffered lines as the paragraph they are. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: breken-ai <312387581+breken-ai@users.noreply.github.com>
|
✅ DCO Check Passed Thanks @breken-ai, all your commits are properly signed off. 🎉 |
Merge Protections🟢 Merge protection satisfied — ready to merge. Show 1 satisfied protection🟢 Enforce conventional commitMake sure that we follow https://www.conventionalcommits.org/en/v1.0.0/
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
ceberam
left a comment
There was a problem hiding this comment.
Thanks @breken-ai for spotting this issue. The bug is real and the diagnosis is accurate.
The fix is correct, well-scoped, and does not regress existing behavior.
Please, address the inline comment and the PR will be ready to be merged.
| # GFM: a table needs a delimiter row right after its header, so | ||
| # lines that merely start with a pipe are a paragraph. Building a | ||
| # table here would drop the second line, taken for that row. |
There was a problem hiding this comment.
The first sentence is useful: it states the GFM rule. The second sentence explains what the old, buggy code did, not what a future maintainer needs to know. It should be removed or rewritten to describe the invariant. For example:
| # GFM: a table needs a delimiter row right after its header, so | |
| # lines that merely start with a pipe are a paragraph. Building a | |
| # table here would drop the second line, taken for that row. | |
| # GFM: a table requires a delimiter row immediately after its header. | |
| # Without one, lines that start with a pipe are ordinary paragraph text. |
|
@breken-ai Please check the PR #4488
I checked and with that approach the case in this PR is no longer an issue, since |
The Markdown backend drops a line of text when a paragraph starts with a pipe but has no GFM delimiter row.
A line starting with
|opens a table in_iterate_elements. When_close_tablebuilds it, it always throws away the second buffered row, assuming it is the|---|---|delimiter row. If there is no delimiter row, that second row is real content and it disappears. The remaining lines turn into a table that was never in the source.This shows up with Pandoc line blocks, which are used for addresses and verse, and with pipe tables written without a delimiter row:
On
main,export_to_markdown()gives:The street line is gone.
| a | b |/| 1 | 2 |/| 3 | 4 |loses1 | 2in the same way.Change: GFM only recognizes a table when a delimiter row follows the header. The pipeless path (
_starts_pipeless_table) already requires one._close_tablenow applies the same rule: if the second buffered row is not a delimiter row, the lines are added as the paragraph they are, so every line is kept. Real tables go through the existing code unchanged.With the fix the example above becomes
| Jane Doe | 221B Baker Street | London NW1 6XEas one text item.Tests:
test_convert_leading_pipe_lines_without_delimiter_row_keep_all_textcovers a delimiter-less pipe table and a pipe-led line followed by plain text. It fails onmain(a table is built and a line is lost) and passes with the fix.tests/test_backend_markdown.py,tests/test_backend_html.py,tests/test_backend_csv.pyandtests/test_backend_asciidoc.pypass (113 passed, 1 skipped) and no reference data changed.ruff check,ruff format --checkandty checkon the touched files are clean.tyshows the same two warnings as onmain.I found and fixed this with help from an AI coding assistant, and I reviewed the change and the test results myself.
Checklist: