CalibRoute is a lightweight Python toolkit for confidence evaluation and
uncertainty-aware routing. It converts model predictions into three actions:
accept, human_review, or abstain.
- Confidence calibration and error-ranking metrics
- Risk-coverage analysis
- Validation-only threshold fitting
- Tie-aware threshold fitting and independent holdout risk assessment
- Batch-size-aware confidence-shift detection
- Financial NER sentence- and entity-level output adapters
- ShiftGuard agent-replay adapter and offline confidence audit
- CSV and JSONL support
- Dependency-free Python API and CLI
Requires Python 3.10 or newer.
Install with python -m pip install calibroute-ai.
To obtain example scripts, clone the repository:
git clone https://github.com/Garyouki/calibroute.git
cd calibrouteUse the installed release for the walkthroughs. For source development, install
the checkout with python -m pip install -e ..
The distribution name is calibroute-ai; the Python import and command are
calibroute. No API key or model service is needed.
Release 0.4.0 adds entity-level NER, ShiftGuard, and comparison reports to
the PyPI package. Upgrade with python -m pip install --upgrade calibroute-ai.
| Capability | PyPI 0.3.0 | PyPI 0.4.0 |
|---|---|---|
audit, fit, validate, route |
Available | Available |
Three-rule comparison report: compare |
Not included | Available |
| Tie-aware thresholds, holdout assessment, and confidence-shift routing | Available | Available |
Financial NER sentence adapter: convert-financial-ner |
Available | Available |
Financial NER entity adapter: convert-financial-ner-entities |
Not included | Available |
ShiftGuard replay adapter: convert-shiftguard |
Not included | Available |
| Document-review walkthrough | Compatible; obtain script from GitHub | Compatible; obtain script from GitHub |
| Financial NER sentence case study | Compatible; obtain script/data from GitHub | Compatible; obtain script/data from GitHub |
| ShiftGuard case study | Not compatible | Compatible; obtain script/data from GitHub |
For reproducible installation, use python -m pip install calibroute-ai==0.4.0.
Example scripts and data are available in the repository; installing the wheel
does not install the example folders. Check out tag v0.4.0 for examples matching
this release; GitHub main may receive further development changes.
To check the active environment, run python -m calibroute.cli --help and
python -m pip show calibroute-ai. Check the listed commands and installation location; record
git rev-parse HEAD when reporting results from a source checkout.
To compare all-accept, a fixed threshold, and CalibRoute on the same labeled batch, see the one-command comparison guide (0.4.0 or newer).
New here? Follow the document-review walkthrough:
one offline script demonstrates development data, a separate holdout, and routing
using the published 0.4.0 package (the script also supports the 0.3.0 API).
For your own exports, follow the CSV onboarding walkthrough
for explicit column mapping, clean installation, and frozen-policy comparison.
The column helper is a repository example added after v0.4.0; obtain it from
main and record the checkout commit. It works with the published 0.4.0 package.
For measured limitations of confidence-shift monitoring, see the synthetic shift study, including false alarms, label-only accuracy losses, and additional review workload.
# Audit labeled predictions
calibroute audit \
--input examples/validation.csv \
--output examples/audit.md
# Fit an acceptance policy on validation data
calibroute fit \
--input examples/validation.csv \
--max-risk 0.20 \
--min-coverage 0.25 \
--risk-method empirical \
--output examples/policy.json
# Route a new prediction batch
calibroute route \
--input examples/production_batch.csv \
--policy examples/policy.json \
--output examples/decisions.csv \
--summary examples/routing-summary.json| Field | Required | Description |
|---|---|---|
id |
recommended | Prediction identifier |
confidence |
yes | Number between 0 and 1 |
correct |
audit and fit only |
Boolean or 1/0 label |
domain |
no | Evaluation slice or deployment domain |
Additional columns are preserved as metadata.
The default fit searches thresholds using pointwise 95% Clopper-Pearson bounds. Threshold selection on the same labels does not provide a 95% guarantee for the selected policy. Freeze the policy, then assess it on a separate IID holdout:
calibroute validate --input holdout.csv --policy examples/policy.json --output holdout-report.jsonExit code 0 means the holdout upper bound meets the declared risk limit; 1 means it does not (including zero accepted samples); 2 means invalid input. Never tune against this holdout or reuse it to select among multiple policies. A pass does not cover distribution shift. See statistical scope.
from calibroute import PredictionRecord, fit_policy, route_batch
validation = [
PredictionRecord("a", 0.98, True),
PredictionRecord("b", 0.82, True),
PredictionRecord("c", 0.55, False),
]
policy = fit_policy(validation, max_risk=0.10, min_coverage=0.50,
risk_method="empirical") # Tiny illustrative sample only.
decisions, summary = route_batch(
[PredictionRecord("new", 0.74)],
policy,
)CalibRoute is an evaluation and routing tool, not a safety certification. Thresholds should be revalidated after changes to the model, task, prompt, or deployment distribution. See design principles for details.
python -m unittest discover -s tests -vSee CONTRIBUTING.md and ROADMAP.md.
Tried it on a task? Share a usage report,
including unsuccessful trials or integration difficulties.
Release preparation is documented in releasing.
The Financial NER case study demonstrates
the adapter and cross-domain failure pattern on 2,098 derived prediction rows.
The ShiftGuard agent example audits 384 frozen
proposals, keeps task templates separate, and reports when no confidence
threshold meets the declared empirical risk target. Use CalibRoute 0.4.0 or newer
for the convert-shiftguard command.
MIT. See LICENSE.