Skip to content

Repository files navigation

Sourcetable Benchmark

The official Sourcetable Benchmark

Results v1

Platform Date all arithmetic connectors files finance formatting formulas styles transform
Sourcetable* 2026-07-03 24/24 1/1 2/2 4/4 4/4 2/2 3/3 3/3 5/5
Microsoft Copilot 2026-07-03 19/24 1/1 0/2 2/4 3/4 2/2 3/3 3/3 5/5
Google Sheets 2026-07-03 17/24 1/1 0/2 2/4 1/4 2/2 3/3 3/3 5/5

* AI mode: Light (even more powerful modes are available)

Summaries

  • Microsoft Copilot: Performs well on typical spreadsheet operations such as data transformations or formulas but does not seem to have full access to the Internet or Code Execution. It does have WebSearch capabilities and can use it to enrich data using more than builtin knowledge but fails on advanced tests that require connecting to third-party services or downloading files.
  • Google Sheets: Also performs well on typical spreadsheet operations such as data transformations or formulas but does not seem to have full access to the Internet (not even basic WebSearch capabilities) or Code Execution. Fails on advanced tests that require connecting to third-party services, downloading files, enriching data with non-builtin knowledge or advanced financial use cases like training and running prediction models or in-depth data analysis on large datasets.

Design Principles

  1. XLSX-In-Out: Each test has at minimum one XLSX input, one XLSX output and one correct XLSX as reference. Additional input, output and reference files are allowed and required for some tests.
  2. Categories: Tests are organized in categories to provide better understanding of what a certain system is good (or bad) at. If a test fits into multiple categories, the best match is used.
  3. Focus: Tests typically focus on a very specific capability only. This is important to ensure that when a certain test fails it's actually because the system lacks this specific capability and not because of any other reason such as not finding the data or not formatting it correctly. Hence, prompts must provide any additional information (format/style/location/...) that is not part of the actually tested capability.
  4. One-Truth: Tests must be completely non-ambiguous and must have exactly one possible solution. This is a requirement for running an automated evaluation of the results and leaving no room for arguing or interpretation. For instance, prompts must exactly specify desired ordering, formatting or anything else that could otherwise easily lead to hundreds of correct permutations of the same result.

Please let us know if you think a test violates these principles.

Running the Benchmark

Please make sure to run the benchmark and evaluation using a tagged version and not the main branch.

Run each test:

  1. Import the XLSX input file (e.g. tests/files-01-input.xlsx)
  2. Upload any other linked input file (e.g. tests/files-01-input.pdf)
  3. Execute the prompt (e.g. tests/files-01-prompt.txt)
  4. Export the result as XLSX (e.g. results-{platform}/files-01.xlsx)
  5. Save additional output artifacts (e.g. results-{platform}/files-01.pdf)

Run the automated evaluation:

  • pip install ironcalc==0.7.0
  • python evaluate.py {platform}

This script compares the files in folder results-{platform} against the correct reference files in folder tests and shows which tests passed or failed as well as the reason for the fail. It will also show the number of passed and failed tests (in total and for each category).

This script evaluates the following conditions:

  • Sheets: The output XLSX must have at least the same sheets as the correct reference. It may have more than that.
  • Bounding Box: The output XLSX must have at least the same cells set as the correct reference. It may have more than that.
  • Type: Each cell that is set in the correct reference must have the same type (e.g. Number) in the output XLSX.
  • Color: Only background color of the cell has to match.
  • Font: Bold/Italic/Underline style has to match, other font attributes (e.g. face or size) are not checked.
  • Formulas: Formulas don't have to match exactly but a cell with a formula in correct reference must have one in output as well. Also, the result of the formula must be identical (see below).
  • Numbers: Numeric values must almost match (epsilon=0.001).
  • Formatted Value: Formatted value has to match (decimal places, thousands separator, date format, ...).

Test Categories

Category Description
arithmetic Basic arithmetic operations such as addition, division etc. (without formulas)
connectors Working with external data sources (like databases) or third-party services
files Working with additional files (attached or downloaded) such as PDF, PNG, DOCX etc.
finance Tests related to stock market, model based predictions or other financial topics
formatting Basic tests related to common Excel formats such as currency, dates etc.
formulas Working with Excel formulas (recognize/create/modify/fix)
styles Testing basic support for Excel styles such as colors, layout etc.
transform All kind of transformations (sorting/filtering/ordering/normalizing/...)

Submitting Results

  • Create a pull request with your files in folder results-{platform}
  • Add your results to the results section of this README.md
  • Provide some sort of proof (e.g. screen recording of your system generating the outputs)

Your system must execute the tests fully autonomously without any further guidance or clarification from the user. You are not allowed to train a model specifically for this test or tailor a custom system prompt for a certain test but instead you must use your default models and system prompts that everyone on your system uses. However, you are allowed to inject a prefix into each test prompt to work around system specific limitations (such as asking for clarification or permissions) as long as it's not giving any further hints about the test case itself. Such an allowed prompt prefix may look like:

THIS IS A FULLY AUTOMATED TEST WITH ID: ${id}
I HEREBY GRANT YOU ALL PERMISSIONS, INCLUDING ANY OPERATION ON ANY SHEET.

IMPORTANT RULES:
- YOU MUST WORK FULLY AUTONOMOUSLY, NO USER CAN RESPOND TO YOU.
- YOU MUST NOT ASK FOR FURTHER CLARIFICATION OR PERMISSION.

VIOLATING ANY OF THESE RULES WILL MAKE YOU FAIL THE TEST.
HERE IS THE TASK:

${prompt}

About

Spreadsheet Benchmark

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages