Skip to content

feat: adding enrichments after the VLM pipeline - #4469

Open
PeterStaar-IBM wants to merge 2 commits into
mainfrom
feat/enable-enrichments-after-vlm-pipeline
Open

PeterStaar-IBM wants to merge 2 commits into
mainfrom
feat/enable-enrichments-after-vlm-pipeline

Conversation

@PeterStaar-IBM

@PeterStaar-IBM PeterStaar-IBM commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Run picture classification, description, and chart extraction after the VLM pipeline. Classification runs before chart extraction.
  • Skip metadata already supplied by the VLM and request only missing chart outputs. Keep picture crops available until enrichment finishes.
  • Restore the older CSV-only Granite Vision Chart2CSV preset and expose chart preset selection in the CLI.
  • Select MLX automatically for Granite Vision 4.1 on compatible Apple Silicon systems, with Transformers fallback. Add an explicit MLX preset and raise the mlx-vlm minimum to 0.7.0.
  • Update model download support, documentation, and regression tests.

Checklist:

  • Documentation has been updated, if necessary.
  • Examples have been added, if necessary.
  • Tests have been added, if necessary.

Signed-off-by: Peter Staar <taa@zurich.ibm.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

✅ DCO Check Passed

Thanks @PeterStaar-IBM, all your commits are properly signed off. 🎉

@mergify

mergify Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Merge Protections

🟢 Merge protection satisfied — ready to merge.

Show 1 satisfied protection

🟢 Enforce conventional commit

Make sure that we follow https://www.conventionalcommits.org/en/v1.0.0/

  • title ~= ^(fix|feat|docs|style|refactor|perf|test|build|ci|chore|revert)(?:\(.+\))?(!)?:

@dosubot

dosubot Bot commented Oct 1, 2026

Copy link
Copy Markdown

📄 Knowledge review

✏️ Suggested updates

1 page suggestion needs review.

Page Library Status
skill.md Docling 🟡 Review
📝 skill.md
@@ -395,6 +395,7 @@
 | `--enrich-picture-classes` | Enable picture classification |
 | `--enrich-picture-description` | Enable picture description |
 | `--enrich-chart-extraction` | Enable chart data extraction |
+| `--chart-extraction-preset PRESET` | Chart extraction preset (granite_vision_v4, granite_vision) |
 | `--image-export-mode MODE` | Image export mode for HTML/MD/JSON outputs |
 | `--num-threads N` | Number of threads |
 | `--device DEVICE` | Accelerator device (cpu, cuda, mps) |
@@ -1114,11 +1115,41 @@
 from docling.datamodel.pipeline_options import (
     PdfPipelineOptions, ChartExtractionModelKind
 )
-
+from docling.datamodel.chart_extraction_options import ChartExtractionVlmEngineOptions
+
+# Using PdfPipelineOptions (for PDF pipeline)
 opts = PdfPipelineOptions(
     do_picture_classification=True,
-    chart_extraction_model=ChartExtractionModelKind.GRANITE_VISION,
-    # or: ChartExtractionModelKind.GRANITE_VISION_V4 (runs chart2csv, chart2code, chart2summary)
+    do_chart_extraction=True,
+    # Default: from_preset('granite_vision_v4') — chart2csv, chart2code, chart2summary
+    chart_extraction_options=ChartExtractionVlmEngineOptions.from_preset('granite_vision_v4'),
+)
+
+# CSV-only model (older Chart2CSV model)
+opts = PdfPipelineOptions(
+    do_picture_classification=True,
+    do_chart_extraction=True,
+    chart_extraction_options=ChartExtractionVlmEngineOptions.from_preset('granite_vision'),
+)
+
+# Backwards-compatible enum (deprecated; use ChartExtractionVlmEngineOptions instead)
+opts = PdfPipelineOptions(
+    do_picture_classification=True,
+    chart_extraction_model=ChartExtractionModelKind.GRANITE_VISION,  # CSV-only model
+    # or: ChartExtractionModelKind.GRANITE_VISION_V4 (chart2csv, chart2code, chart2summary)
+)
+```
+
+Chart extraction enrichments can also be configured for `VlmPipelineOptions` (VLM pipeline):
+
+```python
+from docling.datamodel.pipeline_options import VlmPipelineOptions
+from docling.datamodel.chart_extraction_options import ChartExtractionVlmEngineOptions
+
+# Enable chart extraction in VLM pipeline
+vlm_opts = VlmPipelineOptions(
+    do_chart_extraction=True,
+    chart_extraction_options=ChartExtractionVlmEngineOptions.from_preset('granite_vision_v4'),
 )
 ```
 
@@ -1331,6 +1362,7 @@
 from docling.datamodel.pipeline_options import (
     PdfPipelineOptions,
     PipelineOptions,
+    VlmPipelineOptions,
     # OCR engines
     EasyOcrOptions,
     TesseractCliOcrOptions,
@@ -1348,6 +1380,7 @@
     CodeFormulaVlmOptions,
     ChartExtractionModelKind,
 )
+from docling.datamodel.chart_extraction_options import ChartExtractionVlmEngineOptions
 from docling.datamodel.vlm_engine_options import (
     TransformersVlmEngineOptions,
     MlxVlmEngineOptions,

Accept · Edit · Decline


Leave Feedback Ask Dosu about docling Add Dosu to your team

@codecov

codecov Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 72.28916% with 23 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...g/models/stages/chart_extraction/granite_vision.py 64.86% 9 Missing and 4 partials ⚠️
docling/pipeline/vlm_pipeline.py 63.63% 6 Missing and 2 partials ⚠️
docling/cli/main.py 75.00% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

Signed-off-by: Peter Staar <taa@zurich.ibm.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant