Repository navigation
Routing hygiene for agent harnesses, structured-output fidelity, OpenRouter hardening, shortfall semantic signals - #120
veerareddyvishal144 wants to merge 10 commits into
Conversation
…Router hardening
Found while benchmarking Lynkr tier routing against openrouter/auto on
Terminal-Bench-core 0.1.1 (Terminus agent via LiteLLM). Every change is
gated or scoped so existing clients (Claude Code, Cursor, OpenCode) keep
their current behaviour unless a flag is set.
Request/response fidelity
- Forward the client's output_format / response_format to every provider
(json_schema where the host supports it, json_object otherwise). Lynkr
dropped both, so structured-output clients never got schema enforcement.
- Never run the XML tool-call extractor on a reply that is a JSON object;
it mangled command text containing angle brackets and flipped
stop_reason to tool_use. (databricks.js converter + orchestrator)
- Structured-output requests bypass the response cache: agent harnesses
retry a failed parse with the identical conversation, and a cache hit
returned the same bad reply every time (LYNKR_CACHE_BYPASS_STRUCTURED).
- Remove server-side STANDARD_TOOLS injection from all provider invokers.
A tool-less request stays tool-less; injecting 12 Claude Code tool
schemas turned chat/structured clients into tool-calling agents
(+~2.5k tokens/call, write access to the host workspace). The OpenAI-
compat router's IDE_SAFE_TOOLS injection is left untouched (tied to
CLIENT_TOOL_MAPPINGS for Codex) — follow-up.
- FMT_GUARD_ENABLED=false opt-out for the markdown format guard, whose
"use fenced code blocks" instruction contradicts raw-JSON clients.
Routing on the ask, not the wrapper
- harness-envelope: recognise instruction-schema harness prompts
(Terminus-style preamble + JSON schema + "Instruction:" block) and
expose the instruction as the ask for every turn of the session. The
complexity analyser, intent scorer, agentic detector, kNN query and
the window scorer all use it; later user turns are terminal output and
previously drove force/risk/text scoring (v12: 62 mid-session
COMPLEX->MEDIUM demotions from scoring stdout).
- client-profiles: 'terminus' profile (tool-less by design) + prompt-
pattern detection on the first user message; tool-less profiles are
excluded from side-request detection.
- agentic-detector: drop "solve" from the autonomous regex — harness
boilerplate ("solving command-line tasks") pinned every request to
REASONING.
- FORCE_TIER_PATTERNS=false and RISK_TIER_ESCALATION=false opt-outs for
the keyword escalators (tasks about security/verification are not
themselves risky).
- kNN: LYNKR_KNN_ENABLED=false skips the query and its embedding call;
kNN picks are validated against the configured TIER_* models (the
index remembered a decommissioned model and sent 81 calls to it).
- embeddings: cap input at 5000 chars (nomic-embed-text 2048-token
context), surface Ollama's error body, retry once with halved text on
an input-too-long rejection instead of degrading the provider for 60s.
Provider plumbing and verification
- invokeProvider: [ModelCheck] warning when the served model differs
from the tier's requested model (live incident: TIER zai:glm-5.3-flash
silently served by ZAI_MODEL=glm-5.2 for two full runs).
- Z.AI: unknown tier model ids pass through instead of falling back to
the configured default model.
- OpenRouter: per-model provider pin (OPENROUTER_PROVIDER_ORDER[_MAP]),
per-model reasoning effort, strict json_schema for hosts that honour
it (OPENROUTER_SCHEMA_PROVIDERS), usage.include for provider-reported
cached/reasoning tokens and cost, per-model upstream timeout
(OPENROUTER_TIMEOUT_MS[_MAP]) with failover down the pinned list before
tier-fallback climbs, per-model output cap (OPENROUTER_MAX_TOKENS_MAP),
[ProviderCheck] warning when an unpinned host served. Replies are
converted in the invoker with the hardened converter (cache_read
mapped from prompt_tokens_details.cached_tokens); the orchestrator
passes type:"message" through.
- Fireworks: per-model reasoning effort map, keep thinking enabled for
models that reject thinking:disabled, response_format forwarding.
- OpenAI-compatible: reasoning_effort knob and response_format forwarding.
- Baidu Qianfan: clamp max_tokens to 12288 (hard provider limit).
Tests: npm run test:unit — 1562 pass / 6 fail on both this branch and
main (the 6 are pre-existing, environment-dependent). New flags
documented in .env.example.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ord Jev on the served path
Shortfall built its requirement vector from the STRUCTURAL complexity
dimensions only (message length, tool count, code blocks, turns). On
task-description prompts that vector is nearly constant — measured
0.11–0.34 (median 0.15) across 77 Terminal-Bench tasks, preamble or
not — so cheapest-covering always chose the cheapest model and the Jev
floor vetoed it every time. The main routing path already computes two
semantic difficulty signals shortfall never saw.
- shortfall.liftRequirement(req, { anchorScore, jev }): per head, take
the max of the structural requirement, the anchor-score requirement
(interpolated between adjacent tier profiles at band midpoints 10/35/
63/88) and the Jev requirement (tier-probability-weighted tier
profile). Never lowers a head; null/invalid signals are ignored.
Replay on the 77 tasks: requirement median 0.15 → 0.61, p90 0.30 →
0.81; shortfall now agrees with the anchor+Jev pick on 51/61 (the
remaining 10 are tau permissiveness, not signal).
- routing/index.js feeds analysis.anchorScore / analysis.jev into the
lift and records structuralReq + applied lifts in the shadow log and
shortfallInfo.
- Jev verdicts were null on 100% of served telemetry rows while the
judge was re-tiering ~36% of fresh routes (replay: 11 raised, 17
lowered of 77). Two gaps: the window-scoring wrapper stores the
verdict as `_jev` (telemetry.jevFields only read `jev`/`analysis.jev`),
and the forced-provider hop (api/router.js → invokeModel) never
carried it, so the reconstituted routingResult had no verdict at all.
jevFields now also reads `_jev`; router.js stamps `req.body._jev`;
the forced-path routingResult sets `jev: body._jev`.
- test/shortfall-semantic.test.js (registered in test:unit).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… measured solo runs One command: (1) optionally benchmark each --model solo through a throwaway Lynkr on a spare port (all tiers pinned, own .env/telemetry; needs --run), (2) replay every task's first prompt through the router for its lifted requirement vector (cached), (3) fit per-model/head capability as the highest difficulty level still passed at >= --floor and >= --rel x the best model, tau from the collapse gap, (4) hold-out check of shortfall vs best-model-everywhere on accuracy and cost, report + proposed overrides, --apply to merge into config/model-capabilities.json (backup first). Reuses existing Terminal-Bench run dirs via --model spec=dir. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…d execSync import CI's blocking lint step (eslint --max-warnings 0) failed on two unused bindings: anthropicOutputFormat was added with the structured-output forwarding but the Z.AI path ended up using the OpenAI-compat response_format instead; execSync in cursor-utils.js was unused on main already (main's last two CI runs failed on it). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
🔍 OpenCodeReview found 10 issue(s) in this PR.
|
| fs.copyFileSync(CONFIG_PATH, CONFIG_PATH + `.bak-${Date.now()}`); | ||
| cfg.modelOverrides = { ...(cfg.modelOverrides || {}), ...overrides }; |
There was a problem hiding this comment.
| // absorbs the blip; a provider that is actually down fails twice and | ||
| // degrades exactly as before. | ||
| const TRANSIENT_RETRY_DELAY_MS = 1500; | ||
| const _INPUT_LENGTH_RE = /context length|input length|too long|exceeds? (the )?(maximum|context)|maximum context/i; |
There was a problem hiding this comment.
The _INPUT_LENGTH_RE regex at line 184 is overly broad and could match false positives. The pattern context length|input length|too long|exceeds? (the )?(maximum|context)|maximum context would incorrectly match valid messages containing these words. For example, "too long" appears in many contexts unrelated to errors. Consider making the regex more specific by anchoring to error contexts (e.g., case sensitivity, surrounding keywords like "error", "rejected", or requiring specific error message patterns from known providers).
Suggestion:
| const _INPUT_LENGTH_RE = /context length|input length|too long|exceeds? (the )?(maximum|context)|maximum context/i; | |
| const _INPUT_LENGTH_RE = /\b(context|input)\s+length\s+\b(exceeds?|too long)\b|\bmaximum\s+(context|length)\b|\binput\s+too\s+long\b/i; |
| // fall back to the configured default. Live incident: TIER_*=zai:glm-5.3-flash | ||
| // was silently served by ZAI_MODEL=glm-5.2 for two full benchmark runs | ||
| // (response body "model":"glm-5.2" while Lynkr reported glm-5.3-flash). | ||
| let mappedModel = modelMap[requestedModel] || (body._tierModel ? requestedModel : null) || config.zai.model || "glm-4.7"; |
There was a problem hiding this comment.
At line 2266, body._tierModel is accessed without null checking. When body._tierModel is falsy, the expression (body._tierModel ? requestedModel : null) correctly returns null, but if _tierModel is an empty string or 0, it will use requestedModel (the original value). This could lead to unexpected behavior. Better to explicitly check typeof body._tierModel === 'string' && body._tierModel to ensure it's a valid non-empty string.
| const _bt = block.text.trim(); | ||
| if (_bt.startsWith("{") && _bt.endsWith("}")) continue; |
There was a problem hiding this comment.
In src/orchestrator/index.js, the JSON detection check startsWith("{") && endsWith("}") is fragile and can be bypassed by strings containing JSON anywhere within them. For example, a string like "Here is the JSON: {"key": "value"}" would incorrectly skip extraction. A more robust approach would be to try parsing with JSON.parse() in a try-catch block, or check for a more specific pattern.
Suggestion:
| const _bt = block.text.trim(); | |
| if (_bt.startsWith("{") && _bt.endsWith("}")) continue; | |
| const _bt = block.text.trim(); | |
| try { | |
| JSON.parse(_bt); continue; // Valid JSON object | |
| } catch {} | |
| // Only try to parse first/last char if trim didn't change the string | |
| if (_bt === block.text && _bt.startsWith("{") && _bt.endsWith("}")) continue; |
| if (toolCalls.length === 0 && typeof message.content === "string" && message.content.trim() | ||
| && !(message.content.trim().startsWith("{") && message.content.trim().endsWith("}"))) { |
There was a problem hiding this comment.
In src/orchestrator/index.js, there's duplicated logic for JSON detection on two separate code paths (lines 2318-2319 and 2346-2347). This could be extracted into a helper function like isJsonReply(content) to avoid duplication.
Suggestion:
| if (toolCalls.length === 0 && typeof message.content === "string" && message.content.trim() | |
| && !(message.content.trim().startsWith("{") && message.content.trim().endsWith("}"))) { | |
| const isJsonReply = (text) => { | |
| const trimmed = text.trim(); | |
| if (trimmed === text && trimmed.startsWith("{") && trimmed.endsWith("}")) return true; | |
| try { JSON.parse(trimmed); return true; } catch {} | |
| return false; | |
| }; | |
| // Then use: !isJsonReply(message.content) |
| if (process.env.FORCE_TIER_PATTERNS === "false") return false; | ||
| const content = extractContent(payload); |
There was a problem hiding this comment.
In src/routing/complexity-analyzer.js, both shouldForceCloud and shouldForceReasoning extract content using extractContent(payload) which could return '' (empty string) if payload has no messages. The regex patterns might match unexpected empty content. Consider adding an explicit check for empty content before pattern matching.
Suggestion:
| if (process.env.FORCE_TIER_PATTERNS === "false") return false; | |
| const content = extractContent(payload); | |
| if (process.env.FORCE_TIER_PATTERNS === "false") return false; | |
| const content = extractContent(payload); | |
| if (!content || !content.trim()) return false; |
| if (knnResult && process.env.LYNKR_KNN_ENABLED === 'false') { | ||
| knnResult = null; | ||
| } |
There was a problem hiding this comment.
In src/routing/index.js, the check if (knnResult && process.env.LYNKR_KNN_ENABLED === 'false') is redundant when the outer code already checks config.routing?.knnEnabled !== false && process.env.LYNKR_KNN_ENABLED !== 'false'. The inner check will never be true unless LYNKR_KNN_ENABLED changes between checks (which shouldn't happen). This nested check is defensive but unnecessary.
Suggestion:
| if (knnResult && process.env.LYNKR_KNN_ENABLED === 'false') { | |
| knnResult = null; | |
| } | |
| // This check is redundant given the outer condition already checks LYNKR_KNN_ENABLED | |
| // Consider removing or adding a comment explaining why it's needed |
| const _sel = getModelTierSelector(); | ||
| const _allowed = new Set(); | ||
| for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`); |
There was a problem hiding this comment.
In src/routing/index.js, the kNN tier validation code has duplicated logic in two try-catch blocks (one for success, one for error handling). The same model set construction code appears twice. This could be extracted into a helper function to avoid duplication.
Suggestion:
| const _sel = getModelTierSelector(); | |
| const _allowed = new Set(); | |
| for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`); | |
| const getModelSetFromTiers = () => { | |
| const _sel = getModelTierSelector(); | |
| const _allowed = new Set(); | |
| for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`); | |
| return _allowed; | |
| }; | |
| // Then use: const _allowed = getModelSetFromTiers(); |
| * Lift a structural requirement vector with the semantic signals. | ||
| * @returns {{ req: object, lift: { anchor: object|null, jev: object|null, applied: string[] } }} | ||
| */ | ||
| function liftRequirement(req, { anchorScore = null, jev = null } = {}, tierProfiles = null) { |
There was a problem hiding this comment.
In src/routing/shortfall.js, the liftRequirement function uses default parameters (anchorScore = null, jev = null) which is good. However, it doesn't validate that the input req object actually has the expected properties (HEADS), which could cause issues if an unexpected object is passed.
Suggestion:
| function liftRequirement(req, { anchorScore = null, jev = null } = {}, tierProfiles = null) { | |
| function liftRequirement(req, { anchorScore = null, jev = null } = {}, tierProfiles = null) { | |
| if (!req || typeof req !== 'object') return { req: {}, lift: { anchor: null, jev: null, applied: [] } }; |
| for (const h of HEADS) { | ||
| if (a && a[h] > out[h]) { out[h] = a[h]; if (!applied.includes('anchor')) applied.push('anchor'); } | ||
| if (j && j[h] > out[h]) { out[h] = j[h]; if (!applied.includes('jev')) applied.push('jev'); } |
There was a problem hiding this comment.
The liftRequirement function uses Array.includes() to track which lifts were applied ('anchor', 'jev'). This is inefficient and could be simplified by just counting the number of successful lifts instead of maintaining an array.
Suggestion:
| for (const h of HEADS) { | |
| if (a && a[h] > out[h]) { out[h] = a[h]; if (!applied.includes('anchor')) applied.push('anchor'); } | |
| if (j && j[h] > out[h]) { out[h] = j[h]; if (!applied.includes('jev')) applied.push('jev'); } | |
| const appliedSet = new Set(); | |
| for (const h of HEADS) { | |
| if (a && a[h] > out[h]) { out[h] = a[h]; appliedSet.add('anchor'); } | |
| if (j && j[h] > out[h]) { out[h] = j[h]; appliedSet.add('jev'); } | |
| out[h] = Math.round(Math.max(0, Math.min(1, out[h])) * 1000) / 1000; | |
| } | |
| const applied = [...appliedSet]; |
Removes the "2026-xx-xx local patch" / incident-narration comments added alongside the changes; the PR description carries the rationale. Code is unchanged (eslint clean, unit suite identical: 1569 pass / same 4 pre-existing failures). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| "lint": "eslint src index.js", | ||
| "test": "npm run test:unit && npm run test:performance", | ||
| "test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js", | ||
| "test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js test/shortfall-semantic.test.js", |
| const v = next(); const eq = v.indexOf('='); | ||
| out.models.push(eq > 0 ? { spec: v.slice(0, eq), runDir: expand(v.slice(eq + 1)) } : { spec: v, runDir: null }); |
There was a problem hiding this comment.
Unsanitized file paths: The script reads user-supplied paths (--model=run-dir, --out) and writes to them without validation.
Suggestion:
| const v = next(); const eq = v.indexOf('='); | |
| out.models.push(eq > 0 ? { spec: v.slice(0, eq), runDir: expand(v.slice(eq + 1)) } : { spec: v, runDir: null }); | |
| // Validate runDir exists and is accessible before using | |
| if (m.runDir && !fs.existsSync(m.runDir)) die(`--model run-dir does not exist: ${m.runDir}`); |
| else if (a === '--apply') out.apply = true; | ||
| else if (a === '--dry-run') out.dryRun = true; | ||
| else if (a === '--help' || a === '-h') { console.log(fs.readFileSync(__filename, 'utf8').split('*/')[0].replace(/^\/\*\*?\s?/, '').replace(/^ \* ?/gm, '')); process.exit(0); } | ||
| else die(`unknown arg ${a}`); |
There was a problem hiding this comment.
Inconsistent error messages: --help prints usage before exit, but validation errors via die() don't. Users encountering errors may not know valid arguments.
Suggestion:
| else die(`unknown arg ${a}`); | |
| else { console.error(`unknown arg: ${a}`); console.error('Run with --help for usage'); die(''); } |
| fs.mkdirSync(home, { recursive: true }); fs.mkdirSync(outDir, { recursive: true }); | ||
| // Throwaway Lynkr: copy the operator .env, pin all tiers to this model, | ||
| // own port, own telemetry dir (dotenv + telemetry both key off cwd). | ||
| const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env'); |
There was a problem hiding this comment.
Hardcoded .env paths: The script assumes .env is in process.cwd() or os.homedir(). This may fail in containerized or different working directory contexts.
Suggestion:
| const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env'); | |
| // Add --env-path flag or look for .env relative to ROOT, then cwd, then homedir | |
| const envPaths = [path.join(ROOT, '.env'), path.join(process.cwd(), '.env'), path.join(os.homedir(), '.env')]; |
| fs.mkdirSync(home, { recursive: true }); fs.mkdirSync(outDir, { recursive: true }); | ||
| // Throwaway Lynkr: copy the operator .env, pin all tiers to this model, | ||
| // own port, own telemetry dir (dotenv + telemetry both key off cwd). | ||
| const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env'); |
There was a problem hiding this comment.
Repeated string pattern: fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env') appears twice and should be extracted.
Suggestion:
| const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env'); | |
| function findEnvPath() { | |
| const cwd = path.join(process.cwd(), '.env'); | |
| if (fs.existsSync(cwd)) return cwd; | |
| return path.join(os.homedir(), '.env'); | |
| } |
| detect: { | ||
| headerPatterns: [], | ||
| promptPatterns: [/You are an AI assistant tasked with solving command-line tasks/i], | ||
| minToolFingerprintMatch: 1, | ||
| }, |
| return ''; | ||
| } | ||
| { | ||
| const { harnessAskFromPayload } = require('./harness-envelope'); |
| const HARNESS_PREAMBLE_RES = [ | ||
| /You are an AI assistant tasked with solving command-line tasks/i, | ||
| /"title":\s*"CommandBatchResponse"/, | ||
| ]; |
There was a problem hiding this comment.
The harness prompt detection pattern /You are an AI assistant tasked with solving command-line tasks/i is very broad and may match non-harness prompts from various providers. Consider making it more specific to your actual harness system.
Suggestion:
| const HARNESS_PREAMBLE_RES = [ | |
| /You are an AI assistant tasked with solving command-line tasks/i, | |
| /"title":\s*"CommandBatchResponse"/, | |
| ]; | |
| const HARNESS_PREAMBLE_RES = [ | |
| // Match more specific harness identifiers to avoid false positives | |
| /You are an AI assistant.*?command-line tasks.*?harness/i, | |
| /"title":\s*"CommandBatchResponse"/, | |
| ]; |
| const _allowed = new Set(); | ||
| for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`); |
| const msgs = payload?.messages; | ||
| if (!Array.isArray(msgs)) return { text: null, index: -1 }; | ||
| { | ||
| const { harnessAskFromPayload } = require('./harness-envelope'); |
…on, per-decision effort, route preview/audit, corpus gate, evaluation records, cache probe Declarative routing (config/routing.json, hot-reloaded, validated): - src/routing/routing-config.js — loader/validator; default config has one `legacy` decision (tier from:legacy) so routing is unchanged until edited. Harness recognisers (preamble + instruction regex) live here too; the Terminus pair is the first entry and harness-envelope reads them. - src/routing/signals.js — signal registry: anchor_score (banded), jev, harness, request_field, tool_count, session_turn, session_phase, risk, keyword, legacy_tier, prev_outcome. Pure, fail-open. A condition requires its signal to have matched unless matched:false is explicit. - src/routing/decisions.js — AND/OR/NOT rules, priority, first match wins; tier names or from:legacy|anchor|judge; effort/hosts/plugins per decision; full trace. observe mode records, enforce mode overrides the legacy tier (routing/index.js adopts it and logs a `decision:<name>` escalation). - config/routing.example.json — Terminus-oriented example: regressing sessions escalate, judge/anchor/keyword-hard tasks → COMPLEX@medium pinned to DeepSeek, other harness tasks → MEDIUM@low, structured requests keep caches/format guard off. Turn-outcome attribution (src/routing/outcomes.js): on request N+1 classify turn N as progress / no_progress / regression / provider_error / tool_error / missing from the evidence the request carries (identical-conversation retry, repeated command batch, tool_result is_error, environment error markers) plus the gateway's own record of turn N (telemetry.lastForSession). Only progress/no_progress/regression are model-attributable; rewardFor() returns null for environment noise so it never trains the policy. Streaks skip noise. Per-decision reasoning effort: decisions carry effort; body._effort is honoured by the OpenRouter and Fireworks invokers ahead of the env maps (`none` → reasoning disabled on OpenRouter). Telemetry: decision_name, engine_tier, engine_mode, effort, prev_turn_outcome, prev_turn_attributable columns; engineFields()/lastForSession() helpers; the forced-provider hop carries engine + prev_outcome on the body like _jev. CLI: `lynkr route --preview <req.json> [--json] [--config]` explains every signal, the legacy pick, shortfall's wish and the engine's decision without calling a provider; `lynkr audit <session>|--last N` prints the per-turn ledger (tier, model, decision, effort, latency, cost, judge, previous-turn outcome). Optional X-Lynkr-Decision* response headers (LYNKR_DECISION_HEADERS=true). Corpus gate: test/routing-corpus.test.js replays stored signal snapshots (77 Terminal-Bench first prompts under the example config) through the pure decision step; any rule/config change that alters a decision fails with the task and both decisions printed. scripts/build-routing-corpus.js regenerates. Shortfall: evaluation.records (measured per-model scores) outrank modelOverrides and seeds; calibrate-capabilities --apply writes them. scripts/cache-probe.js: two spaced calls per configured host, reports cached tokens, hit %, cold/warm latency and provider-billed cost (finding: DeepSeek official 97% hits at 1/12 the cold price; Friendli replica-dependent). Tests: routing-decisions (8) + routing-corpus (2) added; test:unit 1577 pass, same 4 pre-existing failures as main; eslint clean on CI scope + new files. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); | ||
| if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); } |
There was a problem hiding this comment.
The audit script reads from a SQLite database file located in the current working directory (.lynkr/telemetry.db) without verifying file permissions or ownership. If the directory has overly permissive permissions, unauthorized users could read sensitive telemetry data including cost, latency, model decisions, and error information.
Suggestion:
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); | |
| if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); } | |
| // Verify database file has appropriate restrictive permissions (e.g., owner-only access) | |
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); | |
| const dbStats = fs.statSync(dbPath); | |
| if (dbStats.mode & 0o444 || dbStats.mode & 0o222) { | |
| console.error('database file has insecure permissions'); | |
| process.exit(2); | |
| } |
| const lastVal = lastIdx >= 0 ? argv[lastIdx + 1] : null; | ||
| const sessionId = argv.find((a) => !a.startsWith('--') && a !== lastVal); | ||
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); | ||
| if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); } |
There was a problem hiding this comment.
Both CLI scripts call process.exit() without returning from the main() function. While this works, it's inconsistent and could cause issues if the functions are called from tests or other modules.
Suggestion:
| if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); } | |
| if (!fs.existsSync(dbPath)) { | |
| console.error(`no telemetry db at ${dbPath}`); | |
| return 2; | |
| } |
| const rows = db.prepare(`SELECT session_id, COUNT(*) turns, MIN(timestamp) t0, MAX(timestamp) t1, SUM(cost_usd) cost, | ||
| SUM(status_code != 200) errors, GROUP_CONCAT(DISTINCT tier) tiers | ||
| FROM routing_telemetry WHERE session_id IS NOT NULL GROUP BY session_id ORDER BY t1 DESC LIMIT ?`).all(n); | ||
| if (json) { console.log(JSON.stringify(rows, null, 2)); return; } |
There was a problem hiding this comment.
JSON.stringify is used for terminal output without any post-processing to sanitize sensitive data that could leak into the output object.
Suggestion:
| if (json) { console.log(JSON.stringify(rows, null, 2)); return; } | |
| // Sanitize or exclude sensitive fields from output | |
| const sanitizedRows = rows.map(r => { | |
| const { ...rest } = r; | |
| // Remove or redact sensitive fields if needed | |
| return rest; | |
| }); | |
| console.log(JSON.stringify(sanitizedRows, null, 2)); |
| const po = r.prev_turn_outcome ? `${r.prev_turn_outcome}${r.prev_turn_attributable ? '' : '(env)'}` : '-'; | ||
| console.log(` ${pad(i + 1, 3)} ${pad(t, 8)} ${pad(r.tier, 9)} ${pad(String(r.model || '').split('/').pop().slice(0, 30), 30)} ${pad(dec.slice(0, 22), 22)} ${pad(r.effort, 6)} ${pad(r.latency_ms, 7)} ${pad(r.status_code, 4)} ${pad((r.cost_usd || 0).toFixed(5), 9)} ${pad(judge, 14)} ${pad(po, 16)}${r.escalation_source ? ' esc:' + r.escalation_source : ''}`); | ||
| }); | ||
| console.log(`\n cost $${cost.toFixed(4)} in ${inTok} (cached ${cached}, ${inTok ? Math.round(100 * cached / (inTok + cached)) : 0}%) out ${outTok} errors ${errs}\n`); |
There was a problem hiding this comment.
The audit script computes cached percentage using integer arithmetic that could produce confusing results for edge cases (e.g., when cached > input tokens).
Suggestion:
| console.log(`\n cost $${cost.toFixed(4)} in ${inTok} (cached ${cached}, ${inTok ? Math.round(100 * cached / (inTok + cached)) : 0}%) out ${outTok} errors ${errs}\n`); | |
| // Add comment explaining the calculation or use clearer logic | |
| console.log(`\n cost $${cost.toFixed(4)} in ${inTok} (cached ${cached}, cached share: ${inTok + cached ? Math.round(100 * cached / (inTok + cached)) : 0}%) out ${outTok} errors ${errs}\n`); |
| console.error('usage: lynkr route --preview <request.json|-> [--json] [--config routing.json]'); | ||
| process.exit(2); | ||
| } | ||
| const raw = file === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(file, 'utf8'); |
There was a problem hiding this comment.
The route script reads request payloads from files or stdin without validating file size limits. A user could provide an extremely large JSON file (e.g., >100MB) that would exhaust available memory when loaded into memory with fs.readFileSync.
Suggestion:
| const raw = file === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(file, 'utf8'); | |
| // Set a reasonable file size limit (e.g., 10MB) to prevent memory exhaustion | |
| const MAX_FILE_SIZE = 10 * 1024 * 1024; | |
| const stats = fs.statSync(file); | |
| if (stats.size > MAX_FILE_SIZE) { | |
| console.error('input file exceeds maximum size limit (10MB)'); | |
| process.exit(2); | |
| } | |
| const raw = fs.readFileSync(file, 'utf8'); |
| function configPath() { | ||
| return process.env.LYNKR_ROUTING_CONFIG ? path.resolve(process.env.LYNKR_ROUTING_CONFIG) : DEFAULT_PATH; | ||
| } |
There was a problem hiding this comment.
The LYNKR_ROUTING_CONFIG environment variable is resolved directly via path.resolve() without sanitization. If this env var is user-controllable, it could enable path traversal attacks. Validate the resolved path stays within the expected directory.
Suggestion:
| function configPath() { | |
| return process.env.LYNKR_ROUTING_CONFIG ? path.resolve(process.env.LYNKR_ROUTING_CONFIG) : DEFAULT_PATH; | |
| } | |
| function configPath() { | |
| const envPath = process.env.LYNKR_ROUTING_CONFIG; | |
| if (!envPath) return DEFAULT_PATH; | |
| const resolved = path.resolve(envPath); | |
| // Validate resolved path stays within allowed directory | |
| const allowedDir = path.resolve(__dirname, '..', '..', 'config'); | |
| if (!resolved.startsWith(allowedDir)) { | |
| logger.warn({ path: resolved }, '[RoutingConfig] invalid config path — using default'); | |
| return DEFAULT_PATH; | |
| } | |
| return resolved; | |
| } |
| for (let i = 0; i < pts.length - 1; i++) { | ||
| const [a, va] = pts[i], [b, vb] = pts[i + 1]; | ||
| if (s >= a && s <= b) { | ||
| const f = (s - a) / (b - a); |
There was a problem hiding this comment.
The interpolation formula (s - a) / (b - a) risks division by zero if two tier midpoints have the same value. Add a guard to handle this edge case before division.
Suggestion:
| for (let i = 0; i < pts.length - 1; i++) { | |
| const [a, va] = pts[i], [b, vb] = pts[i + 1]; | |
| if (s >= a && s <= b) { | |
| const f = (s - a) / (b - a); | |
| for (let i = 0; i < pts.length - 1; i++) { | |
| const [a, va] = pts[i], [b, vb] = pts[i + 1]; | |
| if (s >= a && s <= b) { | |
| const diff = b - a; | |
| const f = diff === 0 ? 0 : (s - a) / diff; |
| for (let i = 0; i < pts.length - 1; i++) { | ||
| const [a, va] = pts[i], [b, vb] = pts[i + 1]; | ||
| if (s >= a && s <= b) { | ||
| const f = (s - a) / (b - a); |
There was a problem hiding this comment.
The interpolation formula (s - a) / (b - a) risks division by zero if two tier midpoints have the same value. Add a guard to handle this edge case before division.
Suggestion:
| for (let i = 0; i < pts.length - 1; i++) { | |
| const [a, va] = pts[i], [b, vb] = pts[i + 1]; | |
| if (s >= a && s <= b) { | |
| const f = (s - a) / (b - a); | |
| for (let i = 0; i < pts.length - 1; i++) { | |
| const [a, va] = pts[i], [b, vb] = pts[i + 1]; | |
| if (s >= a && s <= b) { | |
| const diff = b - a; | |
| const f = diff === 0 ? 0 : (s - a) / diff; |
| harness(ctx) { | ||
| const { harnessAskFromPayload } = require('./harness-envelope'); |
There was a problem hiding this comment.
The harnessAskFromPayload module is required inside two evaluator functions. While Node.js caches modules, this creates unnecessary require calls and could cause circular dependency issues. Move the require to the top of the file.
Suggestion:
| harness(ctx) { | |
| const { harnessAskFromPayload } = require('./harness-envelope'); | |
| const { harnessAskFromPayload } = require('./harness-envelope'); | |
| const EVALUATORS = { | |
| harness(ctx) { |
| harness(ctx) { | ||
| const { harnessAskFromPayload } = require('./harness-envelope'); |
There was a problem hiding this comment.
The harnessAskFromPayload module is required inside two evaluator functions. While Node.js caches modules, this creates unnecessary require calls and could cause circular dependency issues. Move the require to the top of the file.
Suggestion:
| harness(ctx) { | |
| const { harnessAskFromPayload } = require('./harness-envelope'); | |
| const { harnessAskFromPayload } = require('./harness-envelope'); | |
| const EVALUATORS = { | |
| harness(ctx) { |
…check, cascade trigger - src/routing/switch-gate.js: hysteresis over attributable turn outcomes (outcomes.js ring): escalate after N consecutive regressions/no-progress, downgrade after M recoveries above the base tier, min-turns, per-session switch budget, cooldown; environment noise never counts. Wraps checkSessionPin: observe annotates (gate_action/gate_reason), enforce drops the pin and sets a per-session floor the fresh route honours (escalation source switch_gate_floor). - signals.js `complexity`: exemplar contrast — sim(hard centroid) − sim(easy centroid) per head with the configured embedder; exemplars from config/complexity-exemplars.json (scripts/build-complexity-exemplars.js: hard = strong model failed every solo run, easy = cheap model passed). scripts/eval-complexity-signal.js reports AUC vs pass/fail for complexity, anchor, judge, structural and lifted requirement. - src/routing/grounding.js: local NLI (Xenova/nli-deberta-v3-small, int8, CPU) checks completion claims against the last tool/terminal output; verdict supported/unverified/contradicted. Opt-in via LYNKR_GROUNDING (observe|flag); runtime loaded lazily from @huggingface/transformers if present (not a dependency), fails open. Hooked in invokeModel after the provider reply; telemetry grounding_verdict/grounding_contradiction. scripts/validate-grounding.js measures it against a run. Measured on v14: catches 3/28 false completions, wrongly flags 4/39 — no discriminative power on this benchmark (claims describe what the model did and the terminal confirms the command ran; hidden tests fail for other reasons). Kept observe-only. - decisions.js cascade trigger: judge top-two margin < trigger_margin marks the request ambiguous (telemetry cascade_trigger/cascade_margin); enforce path deferred until a verifier proves out. - telemetry: gate_action, gate_reason, cascade_trigger, cascade_margin, grounding_verdict, grounding_contradiction. - config/routing*.json: switch_gate + cascade (observe), complexity signal; example policy uses complexity as a hard-task condition. - test/switch-gate-grounding.test.js (3 tests). test:unit 1582 pass, same 4 pre-existing failures; eslint clean. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…window path The per-message window loop scores a cleaned single message; the harness envelope the `harness` signal keys on is stripped there, so the engine reported `legacy` live while `lynkr route --preview` (full body) matched the harness decisions. The wrapper now evaluates once on the original body and forwards that result. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); } | ||
| const Database = require('better-sqlite3'); | ||
| const db = new Database(dbPath, { readonly: true }); |
There was a problem hiding this comment.
| const sessionId = argv.find((a) => !a.startsWith('--') && a !== lastVal); | ||
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); | ||
| if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); } | ||
| const Database = require('better-sqlite3'); |
There was a problem hiding this comment.
| // Load the operator .env like the server does; silence logs. | ||
| const envPath = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env'); | ||
| require('dotenv').config({ path: envPath }); | ||
| process.env.LOG_LEVEL = 'silent'; process.env.LOG_FILE_ENABLED = 'false'; |
There was a problem hiding this comment.
| "lint": "eslint src index.js", | ||
| "test": "npm run test:unit && npm run test:performance", | ||
| "test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js", | ||
| "test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js test/shortfall-semantic.test.js test/routing-decisions.test.js test/routing-corpus.test.js test/switch-gate-grounding.test.js", |
There was a problem hiding this comment.
New test files are added to the test:unit script: test/shortfall-semantic.test.js, test/routing-decisions.test.js, test/routing-corpus.test.js, and test/switch-gate-grounding.test.js. These correspond to new routing-related source files introduced in the PR. Ensure these test files exist and are properly implemented before merging.
Suggestion:
| "test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js test/shortfall-semantic.test.js test/routing-decisions.test.js test/routing-corpus.test.js test/switch-gate-grounding.test.js", | |
| Verify that the corresponding test files exist in the `test/` directory and contain valid test cases. |
| function findRunRoot(dir) { | ||
| if (fs.existsSync(path.join(dir, 'results.json'))) return dir; | ||
| const subs = fs.readdirSync(dir).filter((d) => fs.existsSync(path.join(dir, d, 'results.json'))).sort().reverse(); | ||
| return subs.length ? path.join(dir, subs[0]) : null; | ||
| } |
There was a problem hiding this comment.
The findRunRoot function doesn't verify that directory entries are actually directories before checking for results.json inside them. If a file with a nested path exists, this could throw an error. Add a directory check to be safe.
Suggestion:
| function findRunRoot(dir) { | |
| if (fs.existsSync(path.join(dir, 'results.json'))) return dir; | |
| const subs = fs.readdirSync(dir).filter((d) => fs.existsSync(path.join(dir, d, 'results.json'))).sort().reverse(); | |
| return subs.length ? path.join(dir, subs[0]) : null; | |
| } | |
| function findRunRoot(dir) { | |
| if (fs.existsSync(path.join(dir, 'results.json'))) return dir; | |
| const subs = fs.readdirSync(dir).filter((d) => fs.statSync(path.join(dir, d)).isDirectory() && fs.existsSync(path.join(dir, d, 'results.json'))).sort().reverse(); | |
| return subs.length ? path.join(dir, subs[0]) : null; | |
| } |
| * Quick check if request should be forced to cloud | ||
| */ | ||
| function shouldForceCloud(payload) { | ||
| if (process.env.FORCE_TIER_PATTERNS === "false") return false; |
| return [{ name: 'harness', preamble: HARNESS_PREAMBLE_RES[0], instruction: HARNESS_INSTRUCTION_RE }]; | ||
| } | ||
|
|
||
| function matchHarness(text) { |
There was a problem hiding this comment.
| const { buildRequirementVector } = require('./capabilities'); | ||
| const sf = require('./shortfall'); | ||
| const req = buildRequirementVector({ dimensions: analysis.breakdown, agenticResult }); | ||
| const _structuralReq = buildRequirementVector({ dimensions: analysis.breakdown, agenticResult }); |
There was a problem hiding this comment.
| const st = fs.statSync(p); | ||
| if (_cache.path === p && _cache.mtime === st.mtimeMs) return _cache.config; | ||
| const raw = JSON.parse(fs.readFileSync(p, 'utf8')); | ||
| const merged = { ...DEFAULT_CONFIG, ...raw, signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) } }; |
There was a problem hiding this comment.
The signal merging on line 120 is shallow: signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) }. This means if you try to override a signal with nested properties (e.g., { anchor: { type: 'anchor_score', extra_option: true } }), you would lose nested properties from DEFAULT_CONFIG. Since the current schema only uses simple { type, ...options } patterns without nested structures, this works today but could break with future schema changes.
Suggestion:
| const merged = { ...DEFAULT_CONFIG, ...raw, signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) } }; | |
| // Consider a deep merge utility if nested signal options are ever needed | |
| const merged = { ...DEFAULT_CONFIG, ...raw, signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) } }; |
| const TTL_MS = 6 * 60 * 60 * 1000; | ||
| const _state = new Map(); // sessionId -> { ts, floor, switches, lastSwitchTurn } |
There was a problem hiding this comment.
Add a _clear() function call at application shutdown or periodically to drain _state entries that have exceeded max_switches_per_session but haven't hit TTL yet. This prevents memory growth from long-lived but inactive sessions.
Suggestion:
| const TTL_MS = 6 * 60 * 60 * 1000; | |
| const _state = new Map(); // sessionId -> { ts, floor, switches, lastSwitchTurn } | |
| const _pinnedProviders = openRouterBody._pinnedProviders; delete openRouterBody._pinnedProviders; | |
| const _schemaOkFn = openRouterBody._schemaOk; delete openRouterBody._schemaOk; | |
| const _timeoutMs = openRouterBody._timeoutMs; delete openRouterBody._timeoutMs; | |
| // ... request code ... | |
| // Restore ALL deleted properties to avoid mutating original body | |
| if (_pinnedProviders) openRouterBody._pinnedProviders = _pinnedProviders; | |
| if (_schemaOkFn) openRouterBody._schemaOk = _schemaOkFn; | |
| if (_timeoutMs) openRouterBody._timeoutMs = _timeoutMs; |
…only in enforce mode or when the engine's tier is the served tier Found during the v15 observe run: effort from a matched decision reached the invoker even though the engine was in observe mode, so tasks where the engine and the legacy chain disagreed ran at the engine's effort on the legacy tier. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| "signals": { | ||
| "anchor": { | ||
| "type": "anchor_score" | ||
| }, |
There was a problem hiding this comment.
The config/routing.json uses camelCase for signal keys (e.g., tool_count, turn, risk) but config/routing.example.json uses snake_case variants. In src/routing/signals.js, signals are accessed using camelCase keys. This inconsistency in the example config could lead to confusion when using it as a template.
| "anchor_bands": { | ||
| "SIMPLE": [ | ||
| 0, | ||
| 20 | ||
| ], |
There was a problem hiding this comment.
| "complexity": { | ||
| "type": "complexity", | ||
| "threshold": 0.6, | ||
| "scale": 0.15 | ||
| } | ||
| }, |
| const requests = []; | ||
| for (const entry of fs.readdirSync(from)) { | ||
| const p = path.join(from, entry); | ||
| if (entry.endsWith('.json') && fs.statSync(p).isFile()) { requests.push({ name: entry.replace(/\.json$/, ''), payload: JSON.parse(fs.readFileSync(p, 'utf8')) }); continue; } |
| function findRunRoot(dir) { | ||
| if (!dir || !fs.existsSync(dir)) return null; | ||
| if (fs.existsSync(path.join(dir, 'results.json'))) return dir; | ||
| const subs = fs.readdirSync(dir).filter((d) => fs.existsSync(path.join(dir, d, 'results.json'))) | ||
| .sort().reverse(); | ||
| return subs.length ? path.join(dir, subs[0]) : null; | ||
| } |
There was a problem hiding this comment.
| if (anchorScore === null || anchorScore === undefined || anchorScore === '') return null; | ||
| const s = Number(anchorScore); | ||
| if (!Number.isFinite(s)) return null; | ||
| const tps = tierProfiles || loadProfiles().tierProfiles; |
There was a problem hiding this comment.
| return { matched: n > 0, value: n, confidence: 1 }; | ||
| }, | ||
| session_turn(ctx) { | ||
| const msgs = Array.isArray(ctx.payload?.messages) ? ctx.payload.messages : []; |
There was a problem hiding this comment.
In session_phase and complexity evaluators, ctx.payload?.messages is accessed without ensuring ctx.payload exists. Although optional chaining (?.) is used, accessing .filter on an undefined result of ctx.payload?.messages could still fail if the property structure is unexpected. Ensure ctx.payload is checked before accessing messages.
| if (!fs.existsSync(file)) return { matched: false, value: null, confidence: null, error: 'no exemplars file' }; | ||
| const st = fs.statSync(file); | ||
| if (!EVALUATORS._cx || EVALUATORS._cx.file !== file || EVALUATORS._cx.mtime !== st.mtimeMs) { | ||
| const doc = JSON.parse(fs.readFileSync(file, 'utf8')); |
| const TIERS = ['SIMPLE', 'MEDIUM', 'COMPLEX', 'REASONING']; | ||
| const PRI = Object.fromEntries(TIERS.map((t, i) => [t, i + 1])); | ||
| const TTL_MS = 6 * 60 * 60 * 1000; | ||
| const _state = new Map(); // sessionId -> { ts, floor, switches, lastSwitchTurn } |
| prev_turn_attributable: po ? (po.attributable ? 1 : 0) : null, | ||
| gate_action: g?.action ?? null, | ||
| gate_reason: g?.reason ?? null, | ||
| cascade_trigger: e?.cascade ? (e.cascade.triggered ? 1 : 0) : null, |
…vLLM Semantic Router Headline table (same VM, concurrency 8, provider-billed cost), setup for all three arms, five findings (auto is one model; both routers lose the same tasks; the gateway decided more than the classifier; Semantic Router lost on timeouts/failover/per-turn classification; structural difficulty scoring is flat), what changed in Lynkr, caveats (noise, effort asymmetry, VM change), and reproduction commands. README gets a one-paragraph pointer next to the RouterArena result. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
| const lastIdx = argv.indexOf('--last'); | ||
| const lastVal = lastIdx >= 0 ? argv[lastIdx + 1] : null; | ||
| const sessionId = argv.find((a) => !a.startsWith('--') && a !== lastVal); | ||
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); |
There was a problem hiding this comment.
Hardcoded database path '.lynkr/telemetry.db' may fail in different deployment environments. Consider adding an environment variable override (e.g., LYNKR_TELEMETRY_DB_PATH) or config option similar to how routing.json is handled.
Suggestion:
| const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); | |
| const dbPath = process.env.LYNKR_TELEMETRY_DB_PATH || path.resolve(process.cwd(), '.lynkr', 'telemetry.db'); |
| const rows = db.prepare(`SELECT session_id, COUNT(*) turns, MIN(timestamp) t0, MAX(timestamp) t1, SUM(cost_usd) cost, | ||
| SUM(status_code != 200) errors, GROUP_CONCAT(DISTINCT tier) tiers | ||
| FROM routing_telemetry WHERE session_id IS NOT NULL GROUP BY session_id ORDER BY t1 DESC LIMIT ?`).all(n); | ||
| if (json) { console.log(JSON.stringify(rows, null, 2)); return; } |
| } | ||
| const raw = file === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(file, 'utf8'); | ||
| let payload; | ||
| try { payload = JSON.parse(raw); } catch (e) { console.error(`not JSON: ${e.message}`); process.exit(2); } |
There was a problem hiding this comment.
JSON.parse() called on user input without structural validation. Consider adding a schema validator to prevent prototype pollution or malformed payloads from causing unexpected behavior.
Suggestion:
| try { payload = JSON.parse(raw); } catch (e) { console.error(`not JSON: ${e.message}`); process.exit(2); } | |
| try { payload = JSON.parse(raw); if (!payload || typeof payload !== 'object') { throw new Error('payload must be an object'); } } catch (e) { console.error(`invalid JSON: ${e.message}`); process.exit(2); } |
| } | ||
|
|
||
| // Load the operator .env like the server does; silence logs. | ||
| const envPath = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env'); |
There was a problem hiding this comment.
Multiple hardcoded .env file paths (cwd then homedir) may be unexpected in containerized environments. Consider adding an environment variable override for consistency with other config paths.
Suggestion:
| const envPath = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env'); | |
| const envPath = process.env.LYNKR_DOTENV_PATH || (fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env')); |
| "prev": { | ||
| "type": "prev_outcome" | ||
| }, |
There was a problem hiding this comment.
| if (a.error || b.error) { console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ERROR ${a.error || b.error}`); continue; } | ||
| const hit = b.prompt ? Math.round(100 * b.cached / b.prompt) : 0; | ||
| const note = b.cached === 0 ? 'NO CACHE HITS REPORTED' : (b.ms < a.ms * 0.8 ? 'warm faster' : 'no latency gain'); | ||
| console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ${String(a.ms).padStart(8)} ${String(b.ms).padStart(8)} ${String(b.cached).padStart(7)} ${String(hit).padStart(5)} ${a.cost != null ? a.cost.toFixed(6).padStart(9) : ' -'} ${b.cost != null ? b.cost.toFixed(6).padStart(9) : ' -'} ${note}`); |
There was a problem hiding this comment.
Inconsistent null comparison: a != null should use strict equality !== null to align with the codebase conventions.
Suggestion:
| console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ${String(a.ms).padStart(8)} ${String(b.ms).padStart(8)} ${String(b.cached).padStart(7)} ${String(hit).padStart(5)} ${a.cost != null ? a.cost.toFixed(6).padStart(9) : ' -'} ${b.cost != null ? b.cost.toFixed(6).padStart(9) : ' -'} ${note}`); | |
| console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ${String(a.ms).padStart(8)} ${String(b.ms).padStart(8)} ${String(b.cached).padStart(7)} ${String(hit).padStart(5)} ${a.cost !== null ? a.cost.toFixed(6).padStart(9) : ' -'} ${b.cost !== null ? b.cost.toFixed(6).padStart(9) : ' -'} ${note}`); |
| const r = spawnSync(tb, ['run', '--dataset', opts.dataset, '--agent', 'terminus', '--model', 'anthropic/claude-sonnet-4-5', | ||
| '--n-concurrent-trials', String(opts.concurrency), '--output-path', outDir], { |
There was a problem hiding this comment.
The lynkr.kill() cleanup only runs if spawnSync completes. If the tb command throws an exception (not a non-zero exit), the orphaned Lynkr process may not be terminated. Consider wrapping the tb spawn in a try-finally to ensure cleanup.
Suggestion:
| const r = spawnSync(tb, ['run', '--dataset', opts.dataset, '--agent', 'terminus', '--model', 'anthropic/claude-sonnet-4-5', | |
| '--n-concurrent-trials', String(opts.concurrency), '--output-path', outDir], { | |
| try { | |
| const r = spawnSync(tb, ['run', '--dataset', opts.dataset, '--agent', 'terminus', '--model', 'anthropic/claude-sonnet-4-5', | |
| '--n-concurrent-trials', String(opts.concurrency), '--output-path', outDir], { | |
| env: { ...process.env, ANTHROPIC_API_BASE: `http://localhost:${opts.port}`, ANTHROPIC_API_KEY: 'lynkr-calibration' }, | |
| stdio: ['ignore', fs.openSync(path.join(outDir, 'tb-stdout.log'), 'a'), fs.openSync(path.join(outDir, 'tb-stdout.log'), 'a')], | |
| maxBuffer: 1 << 26, | |
| }); | |
| // process r.status as before | |
| } finally { | |
| lynkr.kill(); | |
| } |
| function auc(scores, labels) { | ||
| // probability that a random FAIL task scores higher (harder) than a random PASS task | ||
| const pos = [], neg = []; | ||
| scores.forEach((s, i) => { if (s == null) return; (labels[i] ? neg : pos).push(s); }); |
There was a problem hiding this comment.
Inconsistent use of null comparison operators. Line 28 uses == null and != null while the codebase otherwise uses strict equality. This is inconsistent with the strict equality requirement in the code standards.
Suggestion:
| scores.forEach((s, i) => { if (s == null) return; (labels[i] ? neg : pos).push(s); }); | |
| scores.forEach((s, i) => { if (s === null) return; (labels[i] ? neg : pos).push(s); }); |
| for (const p of preds) { | ||
| const s = rows.map((r) => r[p]); | ||
| const a1 = auc(s, rows.map((r) => r.strongPass)); const a2 = cheap ? auc(s, rows.map((r) => r.cheapPass)) : null; | ||
| console.log(`${p.padEnd(12)} ${a1 == null ? ' n/a' : a1.toFixed(3).padStart(15)} ${a2 == null ? ' n/a' : a2.toFixed(3).padStart(10)} ${s.filter((x) => x != null).length}/${rows.length}`); |
There was a problem hiding this comment.
Inconsistent use of null comparison operators in logging output.
Suggestion:
| console.log(`${p.padEnd(12)} ${a1 == null ? ' n/a' : a1.toFixed(3).padStart(15)} ${a2 == null ? ' n/a' : a2.toFixed(3).padStart(10)} ${s.filter((x) => x != null).length}/${rows.length}`); | |
| console.log(`${p.padEnd(12)} ${a1 === null ? ' n/a' : a1.toFixed(3).padStart(15)} ${a2 === null ? ' n/a' : a2.toFixed(3).padStart(10)} ${s.filter((x) => x !== null).length}/${rows.length}`); |
| for (let ai = 0; ai < _orders.length; ai++) { | ||
| const ord = _orders[ai]; | ||
| if (ord.length) { | ||
| openRouterBody.provider = { ...(openRouterBody.provider || {}), order: ord, allow_fallbacks: process.env.OPENROUTER_ALLOW_FALLBACKS === "true" }; | ||
| const wantSchema = !!(_schemaOkFn && _schemaOkFn(ord[0])) || process.env.OPENROUTER_RESPONSE_FORMAT_SCHEMA === "true"; | ||
| const rf = openaiResponseFormat(body, wantSchema); | ||
| if (rf) { if (rf.type === "json_schema") rf.json_schema.strict = true; openRouterBody.response_format = rf; } | ||
| openRouterBody.provider.require_parameters = !!(rf && rf.type === "json_schema"); | ||
| } |
There was a problem hiding this comment.
The openRouterBody.provider object is modified in place during the fallback loop without restoring previous values between iterations. This means that when trying the next provider, the loop reuses provider settings from the previous attempt. The provider should be reset before each iteration to ensure correct behavior.
Summary
Found while benchmarking Lynkr tier routing against
openrouter/autoon Terminal-Bench-core 0.1.1 (Terminus agent via LiteLLM, 80 tasks, ~1,200 calls per run). Every change is gated or scoped so Claude Code, Cursor and OpenCode keep their current behaviour unless a flag is set. Final numbers with these changes: 39/80 for $1.28 vs auto 44/80 for $4.72 on the same VM and evening.Request / response fidelity
output_format/response_formatto every provider (json_schemawhere the host supports it,json_objectotherwise). Lynkr dropped both, so structured-output clients never got schema enforcement.<...>and flippedstop_reasontotool_use.STANDARD_TOOLSinjection from all provider invokers. A tool-less request stays tool-less. Injecting 12 Claude Code tool schemas turned chat / structured clients into tool-calling agents (+~2.5k tokens per call, write access to the host workspace). The OpenAI-compat router'sIDE_SAFE_TOOLSinjection is left untouched (tied toCLIENT_TOOL_MAPPINGSfor Codex) — follow-up.FMT_GUARD_ENABLED=falseopt-out for the markdown format guard (its "use fenced code blocks" instruction contradicts raw-JSON clients).Routing on the ask, not the wrapper
Instruction:block) and expose the instruction as the ask for every turn. Complexity analyser, intent scorer, agentic detector, kNN query and window scorer use it. Later user turns are terminal output and previously drove force / risk / text scoring (62 mid-session COMPLEX→MEDIUM demotions in one run).terminusprofile (tool-less by design) with prompt-pattern detection; tool-less profiles excluded from side-request detection.solvefrom the autonomous regex — "solving command-line tasks" boilerplate pinned every request to REASONING.FORCE_TIER_PATTERNS=false,RISK_TIER_ESCALATION=falseopt-outs for the keyword escalators.LYNKR_KNN_ENABLED=falseskips the query and its embedding call; picks are validated against configuredTIER_*models (the index remembered a decommissioned model and sent 81 calls to it).Provider plumbing and verification
[ModelCheck]warning when the served model differs from the tier's requested model (live incident:TIER_*=zai:glm-5.3-flashsilently served byZAI_MODEL=glm-5.2for two full runs).json_schemafor hosts that honour it,usage.includefor provider-reported cached / reasoning tokens and cost, per-model upstream timeout with failover down the pinned list before tier-fallback climbs, per-model output cap,[ProviderCheck]warning when an unpinned host served. Replies converted in the invoker with cache_read mapped fromprompt_tokens_details.cached_tokens.thinking:disabled,response_formatforwarding.reasoning_effortknob andresponse_formatforwarding.max_tokensto 12288 (hard provider limit).Shortfall: semantic requirement + Jev telemetry (second commit)
shortfall.liftRequirementnow lifts each head with the anchor intent score (interpolated between tier profiles) and the Jev tier probabilities (probability-weighted profile); never lowers. Replay: requirement median 0.15 → 0.61, agreement with the anchor+Jev pick 51/61._jev(reader only knewjev/analysis.jev) and the forced-provider hop never carried it.jevFieldsreads_jev,api/router.jsstampsreq.body._jev, the reconstituted routingResult setsjev. Verified live: verdict, confidence, probabilities and judge model now land on every row.test/shortfall-semantic.test.js.Declarative routing, outcome attribution, operator tooling (third/fourth commits)
config/routing.json, validated, hot-reloaded). Default = onelegacydecision, so nothing changes until edited;config/routing.example.jsonshows a Terminus-oriented policy. AND/OR/NOT rules, priority, per-decision tier/effort/hosts/plugins, observe vs enforce mode, full trace.src/routing/outcomes.js): progress / no_progress / regression / provider_error / tool_error / missing, with a model-attributable flag; environment noise never yields a reward. Recorded per request (prev_turn_outcome).lynkr route --previewexplains every signal and decision for a request without calling a provider;lynkr audit <session>prints the per-turn ledger. OptionalX-Lynkr-Decision*headers.test:unit.model-capabilities.jsonoutrank overrides/seeds;calibrate-capabilities --applywrites them.scripts/cache-probe.jsmeasures prompt-cache behaviour per host.Test plan
npm run test:unit: 1562 pass / 6 fail on this branch and onmain(the 6 are pre-existing, environment-dependent: Jev LRU, TencentDB launcher, downgrade gate, task-ledger).tool_useblock returned).New flags are documented in
.env.example. Happy to split this into smaller PRs (tool-injection removal / structured output / harness awareness / OpenRouter hardening) if preferred.🤖 Generated with Claude Code