Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

GO Transit Reliability Router

Reliability-first routing across the GO Transit network.

Most routing tools optimise for scheduled travel time. This one models real-world reliability — ranking routes by their likelihood of actually working, surfacing active alerts, and explaining tradeoffs in plain language.


Problem

GO Transit services regularly suffer from:

  • Bus no-shows despite showing "on time" in apps
  • Vague service alerts ("operational issues")
  • Transfers with dangerously tight buffers

Solution

GTFS Static  ──► graph (networkx)  ──► Yen's k-shortest paths
GTFS-RT      ──►  reliability score (historical × live modifiers)
                       └──► local LLM (Ollama) ──► plain-language explanation

Routes are generated deterministically. The LLM explains them — it never generates routes or invents transit data. The explanation layer runs locally via Ollama — no API key or cloud account required.


Quickstart

Docker (recommended)

Requires Docker Desktop.

# 1. Configure environment
cp .env.example .env
# Edit .env — minimum required: GTFS_RT_API_KEY (Metrolinx Open Data)

# 2. Start the full stack (PostgreSQL + Ollama + API)
docker compose up -d --build

# 3. Pull the LLM model — one time only (~2 GB, persisted in a volume)
docker compose exec ollama ollama pull llama3.2

# 4. Load GTFS data on first boot (returns 202; runs ~60 s in the background)
curl -X POST http://localhost:8000/ingest/gtfs-static
curl http://localhost:8000/ingest/status   # poll until last_status == "ok"

# 5. Query routes
curl "http://localhost:8000/routes?origin=UN&destination=GL"

# 6. Query with AI explanation
curl "http://localhost:8000/routes?origin=UN&destination=GL&explain=true"

Watch startup logs: docker compose logs -f app

Ollama runs as a bundled service — no host installation needed. The ollama pull step is one-time; the model is stored in the ollama_data Docker volume and survives container restarts. If explanation in the response is a fallback message, check whether the pull completed: docker compose exec ollama ollama list


Local (no Docker)

# 1. Install dependencies
uv sync

# 2. Configure environment
cp .env.example .env

# 3. (Optional) Set up local LLM for ?explain=true
#    Install Ollama from https://ollama.com, then:
ollama pull llama3.2
ollama serve   # OLLAMA_BASE_URL default (localhost:11434) works fine outside Docker

# 4. Start the API (SQLite — no database setup needed)
uv run uvicorn api.main:app --port 8000

# 5. Load GTFS data (first run only; returns 202, ~30 s in the background)
curl -X POST http://localhost:8000/ingest/gtfs-static
curl http://localhost:8000/ingest/status   # poll until last_status == "ok"

# 6. Query routes
curl "http://localhost:8000/routes?origin=UN&destination=GL"

API

GET /routes

Return up to N reliability-scored routes between two stops.

Parameter Required Default Description
origin yes — GTFS stop_id of departure stop
destination yes — GTFS stop_id of arrival stop
departure_time no now Earliest departure as HH:MM or HH:MM:SS. Mutually exclusive with arrive_by
arrive_by no — Latest acceptable arrival as HH:MM or HH:MM:SS, returning the latest-departing options that still make it. GTFS hours apply, so 25:30 is 01:30 next morning
travel_date no today Travel date as YYYY-MM-DD
explain no false true to include LLM explanation (Ollama bundled in Docker; set LLM_PROVIDER in .env)

Responses:

  • 200 — routes found; body contains routes array (+ optional explanation string)

Every leg carries from_lat/from_lon/to_lat/to_lon alongside the stop ids and names, so a map client can draw a route without resolving each intermediate stop through /stops. They are null only if a graph node is missing coordinates, which does not happen for ingested stops.

Times follow the GTFS convention, so a departure or arrival after midnight on the service day reads as 24:00:00 and up — a 25:14:00 departure is 01:14 the following morning, still on the requested travel_date's service day. Clients parsing these as wall-clock times must handle hours past 23.

Each trip leg's risk explains itself: time_bucket names the reliability bucket whose history produced risk_score, and the counters behind it (scheduled_departures, observed_departures, total_delay_seconds, cancellation_count, source) are inlined alongside, so no second request is needed to show why a leg scored as it did. They match the /reliability row for the same (route_id, from_stop_id, time_bucket). When no record exists the counters are zero, source is null, and neutral_prior_used is true.

Trip legs also carry geometry — the stretch of the trip's GTFS shape between that leg's two stops, so the map follows the track rather than drawing a chord between stations. Per leg, so each keeps its own risk colour.

It is a Google encoded polyline at precision 5, roughly 4.4× smaller than the equivalent [lon, lat] arrays. Decode with any standard decoder — @mapbox/polyline's toGeoJSON() returns [lon, lat] pairs ready for MapLibre. (The format is defined latitude-first, so a plain decode() gives [lat, lon].)

The GO feed publishes shapes.txt but no shape_dist_traveled in either shapes.txt or stop_times.txt, so ingest projects each stop onto its shape and stores the resulting index (shape_stop_positions); slicing a leg is then a list slice. Geometry is simplified with Douglas–Peucker at 0.0001° — about 11 m, sub-pixel at city zoom — which takes a typical inter-stop slice from ~120 points to ~28.

geometry is null whenever the trip has no usable shape: no shape_id, no shapes.txt in the feed, or a database ingested before shapes were stored. Clients should fall back to a straight line between the leg's stop coordinates. Walk legs never have it — GTFS publishes no pedestrian geometry.

  • 404 — unknown stop ID, or no routes exist between the stops
  • 422 — invalid parameter format
  • 429 — per-IP rate limit exceeded (see RATE_LIMIT_PER_MINUTE)

Example response:

{
  "routes": [
    {
      "legs": [
        {
          "kind": "trip",
          "from_stop_name": "Union Station GO",
          "to_stop_name": "Bramalea GO",
          "departure_time": "16:22:00",
          "arrival_time": "16:49:00",
          "route_id": "01260426-GT",
          "risk": { "risk_score": 0.2, "risk_label": "Low", "modifiers": [] }
        }
      ],
      "total_travel_seconds": 4860,
      "risk_score": 0.2,
      "risk_label": "Low"
    }
  ]
}

Each leg risk object contains:

Field Type Description
risk_score float 0–1 Combined historical + live risk (higher = riskier)
risk_label string Low (< 0.33) / Medium (< 0.66) / High
modifiers list[str] Human-readable notes (alerts, cancellations, running late, late evening, etc.)
is_cancelled bool true if the trip is currently marked cancelled in GTFS-RT

Same-day trip legs that are currently in the GTFS-RT feed with a non-zero delay also carry live_delay_seconds, expected_departure, and expected_arrival (scheduled + live delay).


GET /stops

Search stops by name substring. Use this to find stop_id values.

Parameter Required Description
query yes Name substring (min 2 characters)

Responses: 200 with array of {stop_id, stop_name, lat, lon, routes_served} objects; empty array if no match.

routes_served is read from the stop_routes table, which ingest materialises from stop_times × trips. Deriving it per request meant a DISTINCT over ~72,000 rows to produce a few dozen pairs (~46 ms); reading it back takes ~5 ms.

curl "http://localhost:8000/stops?query=Guelph"
curl "http://localhost:8000/stops?query=Union"

GET /health

Liveness and data-freshness check.

{
  "status": "ok",
  "timestamp": "2026-02-11T10:00:00",
  "gtfs": {
    "stops": 904,
    "trips": 125245,
    "latest_service_date": "20260601",
    "graph_nodes": 904,
    "graph_edges": 4017,
    "graph_built": true,
    "last_built_at": "2026-02-11T09:30:00",
    "next_refresh_at": "2026-02-12T09:30:00"
  },
  "reliability": {
    "records": 1234,
    "last_seeded_at": "2026-02-11T09:35:00",
    "by_source": {"seed": 1100, "mixed": 130, "observed": 4}
  },
  "gtfs_rt": {
    "polling_active": true,
    "startup_fetch_only": false,
    "last_fetched_at": "2026-02-11T10:00:00+00:00",
    "consecutive_failures": 0,
    "backing_off_until": null,
    "polling_coverage_since": "2026-02-11T09:30:00+00:00",
    "trip_updates": 245,
    "service_alerts": 22,
    "vehicle_positions": 245
  }
}

All counts are 0 (and timestamps null) before /ingest/gtfs-static has been called.


POST /ingest/gtfs-static

Trigger a full GTFS static data refresh and graph rebuild in the background. Runs automatically on a daily schedule; call manually after first install.

If INGEST_API_KEY is set, the request must include X-API-Key: <key>.

Responses: 202 {"status": "accepted", ...} — the ingest (~60 s) runs in the background; poll GET /ingest/status for completion. 401 if the key is wrong/missing; 409 if an ingest is already running.

GET /ingest/status

State of the current or most recent ingest (manual or scheduled): {running, started_at, finished_at, last_status, last_message}. Requires the same optional X-API-Key as the ingest endpoints.

GET /alerts

Active GTFS-RT service alerts (header, description, affected routes/stops, fetched_at) — lets a frontend show a disruption banner without requesting routes. Empty until RT polling is active.


GET /reliability

Inspect the stored counters behind a route's risk score — enough to tell whether a route scores badly from real GTFS-RT observations or from a synthetic prior, which /health (aggregate counts only) cannot.

Parameter Required Default Description
route_id one of these — GTFS route_id
stop_id one of these — GTFS stop_id
time_bucket no — weekday_am_peak, weekday_pm_peak, weekday_offpeak, weekend
limit no 50 Max records, capped at 200

At least one of route_id or stop_id is required and results are capped — this is a lookup for tuning, not a bulk export. score is null with neutral_prior_used: true when a record holds too little data to score, which is exactly where the scorer substitutes the neutral prior.


POST /ingest/reliability-seed

Seed the reliability database from the static GTFS schedule. No GTFS-RT feed required. Uses synthetic per-bucket priors derived from schedule density (see Risk model).

/ingest/gtfs-static already runs a full reseed as part of its pipeline, so calling this manually is only needed to re-seed with a different window_days sample.

Parameter Required Default Description
window_days no 14 Days of schedule to sample (1–90)

Responses: 200 {"status": "ok", "records_written": N, ...} on success; 409 if no GTFS data loaded yet.

# Seed with default 14-day window
curl -X POST http://localhost:8000/ingest/reliability-seed

# Seed using 30 days for a broader sample
curl -X POST "http://localhost:8000/ingest/reliability-seed?window_days=30"

Key stop IDs

/routes accepts any two stop_ids in the ingested feed — 886 stops across 44 routes at the time of writing. These are just the ones used throughout this README, along the Kitchener line.

Stop stop_id
Union Station GO UN
Bloor GO BL
Bramalea GO BE
Brampton Innovation District GO BR
Mount Pleasant GO MO
Georgetown GO GE
Acton GO AC
Guelph Central GO GL
Kitchener GO KI

Risk model

Risk is scored per leg, then the maximum leg risk is used as the route risk (ADR-006 — the weakest link dominates).

Historical prior — per route / stop / time bucket, decayed daily with a 14-day half-life so stats always reflect the recent window:

  • weekday_am_peak (06:00–09:00)
  • weekday_pm_peak (15:00–19:00)
  • weekday_offpeak
  • weekend

Observations come from GTFS-RT: recorded departures with delays, cancellations, and no-shows (scheduled trips that never appeared in any RT feed during continuous polling coverage are swept as misses).

Live modifiers (applied on top of historical prior; per-trip signals apply to same-day travel only):

  • Active service alert for this route or stop
  • Same-day cancellation on this trip
  • Trip currently running late (tiered: ≥5 min, ≥15 min)
  • Missing vehicle position near departure
  • Late-evening departure (after 22:00)

Output: risk_score (0–1) + risk_label (Low / Medium / High)


Configuration

All settings are environment variables (see .env.example):

Variable Default Description
DATABASE_URL sqlite:///data/transit.db SQLite for dev; set to PostgreSQL for prod
GTFS_STATIC_URL Metrolinx CDN URL of GO GTFS ZIP
GTFS_REFRESH_HOURS 24 Interval of the scheduled static refresh
GTFS_RT_API_KEY (blank) Metrolinx Open Data API key; RT polling disabled if unset
GTFS_RT_POLL_SECONDS 30 GTFS-RT poll interval (0 = startup fetch only)
AGENCY_TZ America/Toronto Agency-local timezone for all schedule-time comparisons
LLM_PROVIDER ollama ollama (default, bundled in Docker) or gemini
OLLAMA_BASE_URL http://localhost:11434 Ollama URL; auto-set to http://ollama:11434 inside Docker
OLLAMA_MODEL llama3.2 Ollama model (pull once: docker compose exec ollama ollama pull llama3.2)
GEMINI_API_KEY (blank) Required when LLM_PROVIDER=gemini
GEMINI_MODEL gemini-2.5-flash Gemini model name
INGEST_API_KEY (blank) If set, /ingest/* requires X-API-Key header
RATE_LIMIT_PER_MINUTE 100 Per-IP request cap on /routes, /stops, /alerts (0 disables)
CORS_ORIGINS http://localhost:3000 Comma-separated allowed frontend origins
MAX_ROUTES 5 Max candidate routes returned
MAX_TRANSFERS 2 Hard cap on route changes
MIN_TRANSFER_MINUTES 10 Minimum transfer buffer
MAX_WALK_METRES 500 Walking transfer radius
WALK_SPEED_KPH 4.5 Assumed walking speed for transfer durations

Schema migrations

Schema changes are managed with Alembic.

uv run alembic upgrade head          # apply pending migrations
uv run alembic revision --autogenerate -m "what changed"
uv run alembic downgrade -1          # roll back one

The URL comes from config.DATABASE_URL via alembic/env.py, not from alembic.ini, so migrations always target the same database as the app.

Two things worth knowing:

  • A database created before Alembic already has every table, because init_db() builds the schema with create_all. Mark it current once with uv run alembic stamp head rather than running the baseline migration, which would try to re-create what is already there.
  • Migrations are PostgreSQL-only. Stop.geog is a GeoAlchemy2 Geography column when DATABASE_URL points at PostgreSQL, so the baseline emits geospatial DDL that plain SQLite cannot execute. SQLite (tests, local dev) keeps using init_db()/create_all.

Project structure

transit_planner/
├── api/
│   ├── main.py              FastAPI app assembly (uvicorn entry point)
│   ├── routes.py            Endpoint handlers + route-scoring pipeline
│   ├── lifespan.py          Startup/shutdown, scheduler jobs, ingest slot
│   ├── cache.py             Route cache (TTL, negative, single-flight)
│   ├── ratelimit.py         Per-IP sliding-window rate limiting
│   └── schemas.py           Pydantic response models
├── alembic/
│   ├── env.py               Migration environment (URL from config.DATABASE_URL)
│   └── versions/            Migration scripts
├── db/
│   ├── models.py            SQLAlchemy ORM (GTFS + reliability + shapes)
│   ├── alembic_hooks.py     Autogenerate filter (ignores extension tables)
│   └── session.py           Engine, SessionLocal, get_session
├── graph/
│   └── builder.py           networkx MultiDiGraph construction
├── ingestion/
│   ├── gtfs_static.py       GTFS ZIP download + parse
│   └── gtfs_realtime.py     GTFS-RT protobuf polling (APScheduler)
├── reliability/
│   ├── historical.py        Rolling-window reliability stats
│   └── live.py              Live GTFS-RT risk modifiers
├── llm/
│   └── explainer.py         Local Ollama explanation layer
├── routing/
│   └── engine.py            Yen's k-shortest paths + risk filters
├── config.py                All env-backed configuration
├── alembic.ini              Alembic config (URL comes from env.py)
├── pyproject.toml           Dependencies (managed with uv)
└── .env.example             Environment variable template

Known limitations (v1)

  • Stop-level routing only — no within-stop platform logic.
  • GO buses only — TTC, Brampton Transit, etc. are excluded from routing but their stops may appear in the graph via walk edges.
  • calendar.txt is unused by routing — trips are selected by the GO convention that service_id is a YYYYMMDD date, validated at ingest. A feed switching to standard weekly service_ids would need ServiceCalendar-based resolution.
  • Single uvicorn worker — APScheduler runs in-process; scaling to multiple workers would require moving the scheduler to a separate process.

Data sources

Feed Format Refresh
GO Transit GTFS Static ZIP (CSV) Daily
GTFS-RT Trip Updates Protobuf 30 s
GTFS-RT Vehicle Positions Protobuf 30 s
GTFS-RT Service Alerts Protobuf 30 s

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages