The official Sourcetable Benchmark
| Platform | Date | all | arithmetic | connectors | files | finance | formatting | formulas | styles | transform |
|---|---|---|---|---|---|---|---|---|---|---|
| Sourcetable* | 2026-07-03 |
24/24 |
1/1 |
2/2 |
4/4 |
4/4 |
2/2 |
3/3 |
3/3 |
5/5 |
| Microsoft Copilot | 2026-07-03 |
19/24 |
1/1 |
0/2 |
2/4 |
3/4 |
2/2 |
3/3 |
3/3 |
5/5 |
| Google Sheets | 2026-07-03 |
17/24 |
1/1 |
0/2 |
2/4 |
1/4 |
2/2 |
3/3 |
3/3 |
5/5 |
* AI mode: Light (even more powerful modes are available)
- Microsoft Copilot: Performs well on typical spreadsheet operations such as data transformations or formulas but does not seem to have full access to the Internet or Code Execution. It does have WebSearch capabilities and can use it to enrich data using more than builtin knowledge but fails on advanced tests that require connecting to third-party services or downloading files.
- Google Sheets: Also performs well on typical spreadsheet operations such as data transformations or formulas but does not seem to have full access to the Internet (not even basic WebSearch capabilities) or Code Execution. Fails on advanced tests that require connecting to third-party services, downloading files, enriching data with non-builtin knowledge or advanced financial use cases like training and running prediction models or in-depth data analysis on large datasets.
XLSX-In-Out: Each test has at minimum one XLSX input, one XLSX output and one correct XLSX as reference. Additional input, output and reference files are allowed and required for some tests.Categories: Tests are organized in categories to provide better understanding of what a certain system is good (or bad) at. If a test fits into multiple categories, the best match is used.Focus: Tests typically focus on a very specific capability only. This is important to ensure that when a certain test fails it's actually because the system lacks this specific capability and not because of any other reason such as not finding the data or not formatting it correctly. Hence, prompts must provide any additional information (format/style/location/...) that is not part of the actually tested capability.One-Truth: Tests must be completely non-ambiguous and must have exactly one possible solution. This is a requirement for running an automated evaluation of the results and leaving no room for arguing or interpretation. For instance, prompts must exactly specify desired ordering, formatting or anything else that could otherwise easily lead to hundreds of correct permutations of the same result.
Please let us know if you think a test violates these principles.
Please make sure to run the benchmark and evaluation using a tagged version and not the main branch.
- Import the XLSX input file (e.g.
tests/files-01-input.xlsx) - Upload any other linked input file (e.g.
tests/files-01-input.pdf) - Execute the prompt (e.g.
tests/files-01-prompt.txt) - Export the result as XLSX (e.g.
results-{platform}/files-01.xlsx) - Save additional output artifacts (e.g.
results-{platform}/files-01.pdf)
pip install ironcalc==0.7.0python evaluate.py {platform}
This script compares the files in folder results-{platform} against the correct reference files in folder tests and shows which tests passed or failed as well as the reason for the fail. It will also show the number of passed and failed tests (in total and for each category).
This script evaluates the following conditions:
Sheets: The output XLSX must have at least the same sheets as the correct reference. It may have more than that.Bounding Box: The output XLSX must have at least the same cells set as the correct reference. It may have more than that.Type: Each cell that is set in the correct reference must have the same type (e.g.Number) in the output XLSX.Color: Only background color of the cell has to match.Font: Bold/Italic/Underline style has to match, other font attributes (e.g.faceorsize) are not checked.Formulas: Formulas don't have to match exactly but a cell with a formula in correct reference must have one in output as well. Also, the result of the formula must be identical (see below).Numbers: Numeric values must almost match (epsilon=0.001).Formatted Value: Formatted value has to match (decimal places,thousands separator,date format, ...).
| Category | Description |
|---|---|
| arithmetic | Basic arithmetic operations such as addition, division etc. (without formulas) |
| connectors | Working with external data sources (like databases) or third-party services |
| files | Working with additional files (attached or downloaded) such as PDF, PNG, DOCX etc. |
| finance | Tests related to stock market, model based predictions or other financial topics |
| formatting | Basic tests related to common Excel formats such as currency, dates etc. |
| formulas | Working with Excel formulas (recognize/create/modify/fix) |
| styles | Testing basic support for Excel styles such as colors, layout etc. |
| transform | All kind of transformations (sorting/filtering/ordering/normalizing/...) |
- Create a pull request with your files in folder
results-{platform} - Add your results to the results section of this
README.md - Provide some sort of proof (e.g. screen recording of your system generating the outputs)
Your system must execute the tests fully autonomously without any further guidance or clarification from the user. You are not allowed to train a model specifically for this test or tailor a custom system prompt for a certain test but instead you must use your default models and system prompts that everyone on your system uses. However, you are allowed to inject a prefix into each test prompt to work around system specific limitations (such as asking for clarification or permissions) as long as it's not giving any further hints about the test case itself. Such an allowed prompt prefix may look like:
THIS IS A FULLY AUTOMATED TEST WITH ID: ${id}
I HEREBY GRANT YOU ALL PERMISSIONS, INCLUDING ANY OPERATION ON ANY SHEET.
IMPORTANT RULES:
- YOU MUST WORK FULLY AUTONOMOUSLY, NO USER CAN RESPOND TO YOU.
- YOU MUST NOT ASK FOR FURTHER CLARIFICATION OR PERMISSION.
VIOLATING ANY OF THESE RULES WILL MAKE YOU FAIL THE TEST.
HERE IS THE TASK:
${prompt}