Skip to content

New: Markdown (.md) to Pdf and Docx conversion - #802

Open
Pratham-Solanki911 wants to merge 2 commits into
Tichau:integrationfrom
Pratham-Solanki911:feature/markdown-to-pdf-docx
Open

Pratham-Solanki911 wants to merge 2 commits into
Tichau:integrationfrom
Pratham-Solanki911:feature/markdown-to-pdf-docx

Conversation

@Pratham-Solanki911

Copy link
Copy Markdown

Adds Markdown (.md) input with Pdf and Docx output, plus image output through the existing pdf → image pipeline.

How it works

Markdown is rendered to HTML with Markdig (advanced extensions: tables, task lists, footnotes, etc.). The HTML goes to a temp file, and the existing ConversionJob_Word opens it in Microsoft Word, the same way it opens doc/docx/odt today. From there Word either:

  • exports Pdf with ExportAsFixedFormat (unchanged code path), or
  • saves Docx with SaveAs2(wdFormatXMLDocument) (new).

Reusing Word keeps output consistent with the existing document conversions, gives proper pagination and fonts, and makes the Docx output an editable Word document rather than a flattened one.

Changes

  • New Docx output type, handled by the Word job. This also enables doc/odt → docx for free.
  • .md is registered as a Document input and routed to the Word job.
  • Relative image paths in the markdown resolve against the markdown file's folder, through a <base href> in the generated HTML.
  • Pictures that Word only links when it imports HTML are embedded before saving, so generated docx files are self-contained.
  • The temp HTML file is deleted after conversion, and on failure.
  • Settings.default.xml: md added to "To Pdf", plus a new default "To Docx" preset (md, doc, odt).
  • Installer: Markdig.dll added. Its only dependency, System.Memory, is already shipped.

Installer fix (separate commit)

Product.wxs didn't ship the NetOffice assemblies (NetOffice.dll, OfficeApi.dll, WordApi.dll, ExcelApi.dll, PowerPointApi.dll, VBIDEApi.dll), so every Office-based conversion fails on an installed v2.2. This feature depends on the Word job, so the missing files are added here. It overlaps with the installer part of #779, and whichever PR merges first will make this commit redundant.

Relation to #776

#776 also adds md → pdf, but renders through QuestPDF without Office. This PR takes the Word route, which adds Docx output and matches how the app already handles documents. The trade-off is that it needs Microsoft Word installed, like the other Office conversions.

Testing

I converted one markdown file containing headings, bold/italic/inline code, a link, a table with alignment, a fenced code block, nested and ordered lists, a blockquote, unicode text and a relative image:

  • md → pdf: layout correct, image resolved.
  • md → docx: opens in Word, table and unicode intact, image embedded (r:embed, no r:link).
  • md → png: one page rendered through the existing pdf → image path.
  • No temp HTML left behind.

Markdown files are rendered to html with Markdig, then opened in Microsoft
Word, which exports them like any other Word document. This gives Pdf,
image (via the existing pdf pipeline) and the new Docx output type.

- Add Docx output type, handled by the Word conversion job (also enables
  doc/odt to docx).
- Embed linked pictures so generated docx files are self-contained.
- Resolve relative image paths against the markdown file's folder.
- Add md to the default "To Pdf" preset and a new default "To Docx" preset.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant