Skip to content

Routing hygiene for agent harnesses, structured-output fidelity, OpenRouter hardening, shortfall semantic signals - #120

Open
veerareddyvishal144 wants to merge 10 commits into
mainfrom
bench/routing-and-provider-hardening
Open

veerareddyvishal144 wants to merge 10 commits into
mainfrom
bench/routing-and-provider-hardening

Conversation

@veerareddyvishal144

@veerareddyvishal144 veerareddyvishal144 commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Found while benchmarking Lynkr tier routing against openrouter/auto on Terminal-Bench-core 0.1.1 (Terminus agent via LiteLLM, 80 tasks, ~1,200 calls per run). Every change is gated or scoped so Claude Code, Cursor and OpenCode keep their current behaviour unless a flag is set. Final numbers with these changes: 39/80 for $1.28 vs auto 44/80 for $4.72 on the same VM and evening.

Request / response fidelity

  • Forward the client's output_format / response_format to every provider (json_schema where the host supports it, json_object otherwise). Lynkr dropped both, so structured-output clients never got schema enforcement.
  • Never run the XML tool-call extractor on a JSON-object reply. It mangled command text containing <...> and flipped stop_reason to tool_use.
  • Structured-output requests bypass the response cache. Harnesses retry a failed parse with the identical conversation; a cache hit returned the same bad reply every time.
  • Remove server-side STANDARD_TOOLS injection from all provider invokers. A tool-less request stays tool-less. Injecting 12 Claude Code tool schemas turned chat / structured clients into tool-calling agents (+~2.5k tokens per call, write access to the host workspace). The OpenAI-compat router's IDE_SAFE_TOOLS injection is left untouched (tied to CLIENT_TOOL_MAPPINGS for Codex) — follow-up.
  • FMT_GUARD_ENABLED=false opt-out for the markdown format guard (its "use fenced code blocks" instruction contradicts raw-JSON clients).

Routing on the ask, not the wrapper

  • harness-envelope: recognise instruction-schema harness prompts (fixed preamble + JSON schema + Instruction: block) and expose the instruction as the ask for every turn. Complexity analyser, intent scorer, agentic detector, kNN query and window scorer use it. Later user turns are terminal output and previously drove force / risk / text scoring (62 mid-session COMPLEX→MEDIUM demotions in one run).
  • client-profiles: terminus profile (tool-less by design) with prompt-pattern detection; tool-less profiles excluded from side-request detection.
  • agentic-detector: drop solve from the autonomous regex — "solving command-line tasks" boilerplate pinned every request to REASONING.
  • FORCE_TIER_PATTERNS=false, RISK_TIER_ESCALATION=false opt-outs for the keyword escalators.
  • kNN: LYNKR_KNN_ENABLED=false skips the query and its embedding call; picks are validated against configured TIER_* models (the index remembered a decommissioned model and sent 81 calls to it).
  • embeddings: cap input at 5000 chars (nomic-embed-text's 2048-token context), surface Ollama's error body, retry once with halved text on input-too-long instead of degrading the provider for 60s.

Provider plumbing and verification

  • [ModelCheck] warning when the served model differs from the tier's requested model (live incident: TIER_*=zai:glm-5.3-flash silently served by ZAI_MODEL=glm-5.2 for two full runs).
  • Z.AI: unknown tier model ids pass through instead of falling back to the default model.
  • OpenRouter: per-model provider pin, per-model reasoning effort, strict json_schema for hosts that honour it, usage.include for provider-reported cached / reasoning tokens and cost, per-model upstream timeout with failover down the pinned list before tier-fallback climbs, per-model output cap, [ProviderCheck] warning when an unpinned host served. Replies converted in the invoker with cache_read mapped from prompt_tokens_details.cached_tokens.
  • Fireworks: per-model reasoning effort map, keep thinking enabled for models that reject thinking:disabled, response_format forwarding.
  • OpenAI-compatible: reasoning_effort knob and response_format forwarding.
  • Baidu Qianfan: clamp max_tokens to 12288 (hard provider limit).

Shortfall: semantic requirement + Jev telemetry (second commit)

  • Shortfall's requirement vector came from the structural complexity dimensions only and was nearly constant on task-description prompts (0.11–0.34 across 77 Terminal-Bench tasks), so cheapest-covering always picked the cheapest model. shortfall.liftRequirement now lifts each head with the anchor intent score (interpolated between tier profiles) and the Jev tier probabilities (probability-weighted profile); never lowers. Replay: requirement median 0.15 → 0.61, agreement with the anchor+Jev pick 51/61.
  • Jev verdicts were null on 100% of served telemetry rows while the judge re-tiered ~36% of fresh routes. Two gaps: the window wrapper stores the verdict as _jev (reader only knew jev/analysis.jev) and the forced-provider hop never carried it. jevFields reads _jev, api/router.js stamps req.body._jev, the reconstituted routingResult sets jev. Verified live: verdict, confidence, probabilities and judge model now land on every row.
  • New test/shortfall-semantic.test.js.

Declarative routing, outcome attribution, operator tooling (third/fourth commits)

  • Signals → decisions as config (config/routing.json, validated, hot-reloaded). Default = one legacy decision, so nothing changes until edited; config/routing.example.json shows a Terminus-oriented policy. AND/OR/NOT rules, priority, per-decision tier/effort/hosts/plugins, observe vs enforce mode, full trace.
  • Turn-outcome attribution (src/routing/outcomes.js): progress / no_progress / regression / provider_error / tool_error / missing, with a model-attributable flag; environment noise never yields a reward. Recorded per request (prev_turn_outcome).
  • Per-decision reasoning effort honoured by the OpenRouter/Fireworks invokers.
  • lynkr route --preview explains every signal and decision for a request without calling a provider; lynkr audit <session> prints the per-turn ledger. Optional X-Lynkr-Decision* headers.
  • Corpus regression gate: 77 signal-snapshot fixtures replayed through the pure decision step in test:unit.
  • Evaluation records in model-capabilities.json outrank overrides/seeds; calibrate-capabilities --apply writes them. scripts/cache-probe.js measures prompt-cache behaviour per host.

Test plan

  • npm run test:unit: 1562 pass / 6 fail on this branch and on main (the 6 are pre-existing, environment-dependent: Jev LRU, TencentDB launcher, downgrade gate, task-ledger).
  • Live: Terminus harness end-to-end through Lynkr on OpenRouter (Friendli glm-5.3-flash cheap tier, DeepSeek official strong tier): two full 80-task runs, 0 wrong-model warnings, 0 harness parse failures, failover drill (forced timeout → next host → tier climb) verified.
  • Requests that carry their own tools still forward them (probed: tool_use block returned).

New flags are documented in .env.example. Happy to split this into smaller PRs (tool-injection removal / structured output / harness awareness / OpenRouter hardening) if preferred.

🤖 Generated with Claude Code

veerareddyvishal144 and others added 4 commits October 3, 2026 21:43
…Router hardening

Found while benchmarking Lynkr tier routing against openrouter/auto on
Terminal-Bench-core 0.1.1 (Terminus agent via LiteLLM). Every change is
gated or scoped so existing clients (Claude Code, Cursor, OpenCode) keep
their current behaviour unless a flag is set.

Request/response fidelity
- Forward the client's output_format / response_format to every provider
  (json_schema where the host supports it, json_object otherwise). Lynkr
  dropped both, so structured-output clients never got schema enforcement.
- Never run the XML tool-call extractor on a reply that is a JSON object;
  it mangled command text containing angle brackets and flipped
  stop_reason to tool_use. (databricks.js converter + orchestrator)
- Structured-output requests bypass the response cache: agent harnesses
  retry a failed parse with the identical conversation, and a cache hit
  returned the same bad reply every time (LYNKR_CACHE_BYPASS_STRUCTURED).
- Remove server-side STANDARD_TOOLS injection from all provider invokers.
  A tool-less request stays tool-less; injecting 12 Claude Code tool
  schemas turned chat/structured clients into tool-calling agents
  (+~2.5k tokens/call, write access to the host workspace). The OpenAI-
  compat router's IDE_SAFE_TOOLS injection is left untouched (tied to
  CLIENT_TOOL_MAPPINGS for Codex) — follow-up.
- FMT_GUARD_ENABLED=false opt-out for the markdown format guard, whose
  "use fenced code blocks" instruction contradicts raw-JSON clients.

Routing on the ask, not the wrapper
- harness-envelope: recognise instruction-schema harness prompts
  (Terminus-style preamble + JSON schema + "Instruction:" block) and
  expose the instruction as the ask for every turn of the session. The
  complexity analyser, intent scorer, agentic detector, kNN query and
  the window scorer all use it; later user turns are terminal output and
  previously drove force/risk/text scoring (v12: 62 mid-session
  COMPLEX->MEDIUM demotions from scoring stdout).
- client-profiles: 'terminus' profile (tool-less by design) + prompt-
  pattern detection on the first user message; tool-less profiles are
  excluded from side-request detection.
- agentic-detector: drop "solve" from the autonomous regex — harness
  boilerplate ("solving command-line tasks") pinned every request to
  REASONING.
- FORCE_TIER_PATTERNS=false and RISK_TIER_ESCALATION=false opt-outs for
  the keyword escalators (tasks about security/verification are not
  themselves risky).
- kNN: LYNKR_KNN_ENABLED=false skips the query and its embedding call;
  kNN picks are validated against the configured TIER_* models (the
  index remembered a decommissioned model and sent 81 calls to it).
- embeddings: cap input at 5000 chars (nomic-embed-text 2048-token
  context), surface Ollama's error body, retry once with halved text on
  an input-too-long rejection instead of degrading the provider for 60s.

Provider plumbing and verification
- invokeProvider: [ModelCheck] warning when the served model differs
  from the tier's requested model (live incident: TIER zai:glm-5.3-flash
  silently served by ZAI_MODEL=glm-5.2 for two full runs).
- Z.AI: unknown tier model ids pass through instead of falling back to
  the configured default model.
- OpenRouter: per-model provider pin (OPENROUTER_PROVIDER_ORDER[_MAP]),
  per-model reasoning effort, strict json_schema for hosts that honour
  it (OPENROUTER_SCHEMA_PROVIDERS), usage.include for provider-reported
  cached/reasoning tokens and cost, per-model upstream timeout
  (OPENROUTER_TIMEOUT_MS[_MAP]) with failover down the pinned list before
  tier-fallback climbs, per-model output cap (OPENROUTER_MAX_TOKENS_MAP),
  [ProviderCheck] warning when an unpinned host served. Replies are
  converted in the invoker with the hardened converter (cache_read
  mapped from prompt_tokens_details.cached_tokens); the orchestrator
  passes type:"message" through.
- Fireworks: per-model reasoning effort map, keep thinking enabled for
  models that reject thinking:disabled, response_format forwarding.
- OpenAI-compatible: reasoning_effort knob and response_format forwarding.
- Baidu Qianfan: clamp max_tokens to 12288 (hard provider limit).

Tests: npm run test:unit — 1562 pass / 6 fail on both this branch and
main (the 6 are pre-existing, environment-dependent). New flags
documented in .env.example.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ord Jev on the served path

Shortfall built its requirement vector from the STRUCTURAL complexity
dimensions only (message length, tool count, code blocks, turns). On
task-description prompts that vector is nearly constant — measured
0.11–0.34 (median 0.15) across 77 Terminal-Bench tasks, preamble or
not — so cheapest-covering always chose the cheapest model and the Jev
floor vetoed it every time. The main routing path already computes two
semantic difficulty signals shortfall never saw.

- shortfall.liftRequirement(req, { anchorScore, jev }): per head, take
  the max of the structural requirement, the anchor-score requirement
  (interpolated between adjacent tier profiles at band midpoints 10/35/
  63/88) and the Jev requirement (tier-probability-weighted tier
  profile). Never lowers a head; null/invalid signals are ignored.
  Replay on the 77 tasks: requirement median 0.15 → 0.61, p90 0.30 →
  0.81; shortfall now agrees with the anchor+Jev pick on 51/61 (the
  remaining 10 are tau permissiveness, not signal).
- routing/index.js feeds analysis.anchorScore / analysis.jev into the
  lift and records structuralReq + applied lifts in the shadow log and
  shortfallInfo.
- Jev verdicts were null on 100% of served telemetry rows while the
  judge was re-tiering ~36% of fresh routes (replay: 11 raised, 17
  lowered of 77). Two gaps: the window-scoring wrapper stores the
  verdict as `_jev` (telemetry.jevFields only read `jev`/`analysis.jev`),
  and the forced-provider hop (api/router.js → invokeModel) never
  carried it, so the reconstituted routingResult had no verdict at all.
  jevFields now also reads `_jev`; router.js stamps `req.body._jev`;
  the forced-path routingResult sets `jev: body._jev`.
- test/shortfall-semantic.test.js (registered in test:unit).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… measured solo runs

One command: (1) optionally benchmark each --model solo through a throwaway
Lynkr on a spare port (all tiers pinned, own .env/telemetry; needs --run),
(2) replay every task's first prompt through the router for its lifted
requirement vector (cached), (3) fit per-model/head capability as the
highest difficulty level still passed at >= --floor and >= --rel x the
best model, tau from the collapse gap, (4) hold-out check of shortfall vs
best-model-everywhere on accuracy and cost, report + proposed overrides,
--apply to merge into config/model-capabilities.json (backup first).
Reuses existing Terminal-Bench run dirs via --model spec=dir.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…d execSync import

CI's blocking lint step (eslint --max-warnings 0) failed on two unused
bindings: anthropicOutputFormat was added with the structured-output
forwarding but the Z.AI path ended up using the OpenAI-compat
response_format instead; execSync in cursor-utils.js was unused on main
already (main's last two CI runs failed on it).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

🔍 OpenCodeReview found 10 issue(s) in this PR.

  • ✅ Successfully posted inline: 10 comment(s)

Comment on lines +415 to +416
fs.copyFileSync(CONFIG_PATH, CONFIG_PATH + `.bak-${Date.now()}`);
cfg.modelOverrides = { ...(cfg.modelOverrides || {}), ...overrides };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · medium
The backup file write is not validated before overwriting the config. If copyFileSync fails (e.g., permission issues), the config could be partially overwritten or lost. Add a check after copyFileSync to ensure the backup exists before proceeding with the rewrite.

Comment thread src/cache/embeddings.js
// absorbs the blip; a provider that is actually down fails twice and
// degrades exactly as before.
const TRANSIENT_RETRY_DELAY_MS = 1500;
const _INPUT_LENGTH_RE = /context length|input length|too long|exceeds? (the )?(maximum|context)|maximum context/i;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · low
The _INPUT_LENGTH_RE regex at line 184 is overly broad and could match false positives. The pattern context length|input length|too long|exceeds? (the )?(maximum|context)|maximum context would incorrectly match valid messages containing these words. For example, "too long" appears in many contexts unrelated to errors. Consider making the regex more specific by anchoring to error contexts (e.g., case sensitivity, surrounding keywords like "error", "rejected", or requiring specific error message patterns from known providers).

Suggestion:

Suggested change
const _INPUT_LENGTH_RE = /context length|input length|too long|exceeds? (the )?(maximum|context)|maximum context/i;
const _INPUT_LENGTH_RE = /\b(context|input)\s+length\s+\b(exceeds?|too long)\b|\bmaximum\s+(context|length)\b|\binput\s+too\s+long\b/i;

Comment thread src/clients/databricks.js
// fall back to the configured default. Live incident: TIER_*=zai:glm-5.3-flash
// was silently served by ZAI_MODEL=glm-5.2 for two full benchmark runs
// (response body "model":"glm-5.2" while Lynkr reported glm-5.3-flash).
let mappedModel = modelMap[requestedModel] || (body._tierModel ? requestedModel : null) || config.zai.model || "glm-4.7";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · medium
At line 2266, body._tierModel is accessed without null checking. When body._tierModel is falsy, the expression (body._tierModel ? requestedModel : null) correctly returns null, but if _tierModel is an empty string or 0, it will use requestedModel (the original value). This could lead to unexpected behavior. Better to explicitly check typeof body._tierModel === 'string' && body._tierModel to ensure it's a valid non-empty string.

Comment thread src/orchestrator/index.js
Comment on lines +2318 to +2319
const _bt = block.text.trim();
if (_bt.startsWith("{") && _bt.endsWith("}")) continue;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · medium
In src/orchestrator/index.js, the JSON detection check startsWith("{") && endsWith("}") is fragile and can be bypassed by strings containing JSON anywhere within them. For example, a string like "Here is the JSON: {"key": "value"}" would incorrectly skip extraction. A more robust approach would be to try parsing with JSON.parse() in a try-catch block, or check for a more specific pattern.

Suggestion:

Suggested change
const _bt = block.text.trim();
if (_bt.startsWith("{") && _bt.endsWith("}")) continue;
const _bt = block.text.trim();
try {
JSON.parse(_bt); continue; // Valid JSON object
} catch {}
// Only try to parse first/last char if trim didn't change the string
if (_bt === block.text && _bt.startsWith("{") && _bt.endsWith("}")) continue;

Comment thread src/orchestrator/index.js
Comment on lines +2346 to +2347
if (toolCalls.length === 0 && typeof message.content === "string" && message.content.trim()
&& !(message.content.trim().startsWith("{") && message.content.trim().endsWith("}"))) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
In src/orchestrator/index.js, there's duplicated logic for JSON detection on two separate code paths (lines 2318-2319 and 2346-2347). This could be extracted into a helper function like isJsonReply(content) to avoid duplication.

Suggestion:

Suggested change
if (toolCalls.length === 0 && typeof message.content === "string" && message.content.trim()
&& !(message.content.trim().startsWith("{") && message.content.trim().endsWith("}"))) {
const isJsonReply = (text) => {
const trimmed = text.trim();
if (trimmed === text && trimmed.startsWith("{") && trimmed.endsWith("}")) return true;
try { JSON.parse(trimmed); return true; } catch {}
return false;
};
// Then use: !isJsonReply(message.content)

Comment on lines +997 to 998
if (process.env.FORCE_TIER_PATTERNS === "false") return false;
const content = extractContent(payload);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · low
In src/routing/complexity-analyzer.js, both shouldForceCloud and shouldForceReasoning extract content using extractContent(payload) which could return '' (empty string) if payload has no messages. The regex patterns might match unexpected empty content. Consider adding an explicit check for empty content before pattern matching.

Suggestion:

Suggested change
if (process.env.FORCE_TIER_PATTERNS === "false") return false;
const content = extractContent(payload);
if (process.env.FORCE_TIER_PATTERNS === "false") return false;
const content = extractContent(payload);
if (!content || !content.trim()) return false;

Comment thread src/routing/index.js
Comment on lines +1774 to +1776
if (knnResult && process.env.LYNKR_KNN_ENABLED === 'false') {
knnResult = null;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
In src/routing/index.js, the check if (knnResult && process.env.LYNKR_KNN_ENABLED === 'false') is redundant when the outer code already checks config.routing?.knnEnabled !== false && process.env.LYNKR_KNN_ENABLED !== 'false'. The inner check will never be true unless LYNKR_KNN_ENABLED changes between checks (which shouldn't happen). This nested check is defensive but unnecessary.

Suggestion:

Suggested change
if (knnResult && process.env.LYNKR_KNN_ENABLED === 'false') {
knnResult = null;
}
// This check is redundant given the outer condition already checks LYNKR_KNN_ENABLED
// Consider removing or adding a comment explaining why it's needed

Comment thread src/routing/index.js
Comment on lines +1779 to +1781
const _sel = getModelTierSelector();
const _allowed = new Set();
for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
In src/routing/index.js, the kNN tier validation code has duplicated logic in two try-catch blocks (one for success, one for error handling). The same model set construction code appears twice. This could be extracted into a helper function to avoid duplication.

Suggestion:

Suggested change
const _sel = getModelTierSelector();
const _allowed = new Set();
for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`);
const getModelSetFromTiers = () => {
const _sel = getModelTierSelector();
const _allowed = new Set();
for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`);
return _allowed;
};
// Then use: const _allowed = getModelSetFromTiers();

Comment thread src/routing/shortfall.js
* Lift a structural requirement vector with the semantic signals.
* @returns {{ req: object, lift: { anchor: object|null, jev: object|null, applied: string[] } }}
*/
function liftRequirement(req, { anchorScore = null, jev = null } = {}, tierProfiles = null) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · low
In src/routing/shortfall.js, the liftRequirement function uses default parameters (anchorScore = null, jev = null) which is good. However, it doesn't validate that the input req object actually has the expected properties (HEADS), which could cause issues if an unexpected object is passed.

Suggestion:

Suggested change
function liftRequirement(req, { anchorScore = null, jev = null } = {}, tierProfiles = null) {
function liftRequirement(req, { anchorScore = null, jev = null } = {}, tierProfiles = null) {
if (!req || typeof req !== 'object') return { req: {}, lift: { anchor: null, jev: null, applied: [] } };

Comment thread src/routing/shortfall.js
Comment on lines +366 to +368
for (const h of HEADS) {
if (a && a[h] > out[h]) { out[h] = a[h]; if (!applied.includes('anchor')) applied.push('anchor'); }
if (j && j[h] > out[h]) { out[h] = j[h]; if (!applied.includes('jev')) applied.push('jev'); }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

performance · low
The liftRequirement function uses Array.includes() to track which lifts were applied ('anchor', 'jev'). This is inefficient and could be simplified by just counting the number of successful lifts instead of maintaining an array.

Suggestion:

Suggested change
for (const h of HEADS) {
if (a && a[h] > out[h]) { out[h] = a[h]; if (!applied.includes('anchor')) applied.push('anchor'); }
if (j && j[h] > out[h]) { out[h] = j[h]; if (!applied.includes('jev')) applied.push('jev'); }
const appliedSet = new Set();
for (const h of HEADS) {
if (a && a[h] > out[h]) { out[h] = a[h]; appliedSet.add('anchor'); }
if (j && j[h] > out[h]) { out[h] = j[h]; appliedSet.add('jev'); }
out[h] = Math.round(Math.max(0, Math.min(1, out[h])) * 1000) / 1000;
}
const applied = [...appliedSet];

Removes the "2026-xx-xx local patch" / incident-narration comments added
alongside the changes; the PR description carries the rationale. Code is
unchanged (eslint clean, unit suite identical: 1569 pass / same 4
pre-existing failures).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment thread package.json Outdated
"lint": "eslint src index.js",
"test": "npm run test:unit && npm run test:performance",
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js",
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js test/shortfall-semantic.test.js",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
package.json update: Added test/shortfall-semantic.test.js to test:unit. No dependency conflicts or duplicate declarations found. All tools referenced in scripts (eslint, pino-pretty) are declared in devDependencies.

Comment on lines +75 to +76
const v = next(); const eq = v.indexOf('=');
out.models.push(eq > 0 ? { spec: v.slice(0, eq), runDir: expand(v.slice(eq + 1)) } : { spec: v, runDir: null });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · low
Unsanitized file paths: The script reads user-supplied paths (--model=run-dir, --out) and writes to them without validation.

Suggestion:

Suggested change
const v = next(); const eq = v.indexOf('=');
out.models.push(eq > 0 ? { spec: v.slice(0, eq), runDir: expand(v.slice(eq + 1)) } : { spec: v, runDir: null });
// Validate runDir exists and is accessible before using
if (m.runDir && !fs.existsSync(m.runDir)) die(`--model run-dir does not exist: ${m.runDir}`);

else if (a === '--apply') out.apply = true;
else if (a === '--dry-run') out.dryRun = true;
else if (a === '--help' || a === '-h') { console.log(fs.readFileSync(__filename, 'utf8').split('*/')[0].replace(/^\/\*\*?\s?/, '').replace(/^ \* ?/gm, '')); process.exit(0); }
else die(`unknown arg ${a}`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · low
Inconsistent error messages: --help prints usage before exit, but validation errors via die() don't. Users encountering errors may not know valid arguments.

Suggestion:

Suggested change
else die(`unknown arg ${a}`);
else { console.error(`unknown arg: ${a}`); console.error('Run with --help for usage'); die(''); }

fs.mkdirSync(home, { recursive: true }); fs.mkdirSync(outDir, { recursive: true });
// Throwaway Lynkr: copy the operator .env, pin all tiers to this model,
// own port, own telemetry dir (dotenv + telemetry both key off cwd).
const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Hardcoded .env paths: The script assumes .env is in process.cwd() or os.homedir(). This may fail in containerized or different working directory contexts.

Suggestion:

Suggested change
const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env');
// Add --env-path flag or look for .env relative to ROOT, then cwd, then homedir
const envPaths = [path.join(ROOT, '.env'), path.join(process.cwd(), '.env'), path.join(os.homedir(), '.env')];

fs.mkdirSync(home, { recursive: true }); fs.mkdirSync(outDir, { recursive: true });
// Throwaway Lynkr: copy the operator .env, pin all tiers to this model,
// own port, own telemetry dir (dotenv + telemetry both key off cwd).
const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Repeated string pattern: fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env') appears twice and should be extracted.

Suggestion:

Suggested change
const srcEnv = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(os.homedir(), '.env');
function findEnvPath() {
const cwd = path.join(process.cwd(), '.env');
if (fs.existsSync(cwd)) return cwd;
return path.join(os.homedir(), '.env');
}

Comment on lines +41 to +45
detect: {
headerPatterns: [],
promptPatterns: [/You are an AI assistant tasked with solving command-line tasks/i],
minToolFingerprintMatch: 1,
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The terminus profile's prompt pattern matches the HARNESS_PREAMBLE_RES pattern in harness-envelope.js. This creates tight coupling between client detection and harness parsing logic. Consider unifying these patterns or moving to a shared configuration file.

return '';
}
{
const { harnessAskFromPayload } = require('./harness-envelope');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
Same dynamic require pattern here - could cause circular dependency issues if the module loading order changes.

Comment on lines +51 to +54
const HARNESS_PREAMBLE_RES = [
/You are an AI assistant tasked with solving command-line tasks/i,
/"title":\s*"CommandBatchResponse"/,
];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · high
The harness prompt detection pattern /You are an AI assistant tasked with solving command-line tasks/i is very broad and may match non-harness prompts from various providers. Consider making it more specific to your actual harness system.

Suggestion:

Suggested change
const HARNESS_PREAMBLE_RES = [
/You are an AI assistant tasked with solving command-line tasks/i,
/"title":\s*"CommandBatchResponse"/,
];
const HARNESS_PREAMBLE_RES = [
// Match more specific harness identifiers to avoid false positives
/You are an AI assistant.*?command-line tasks.*?harness/i,
/"title":\s*"CommandBatchResponse"/,
];

Comment thread src/routing/index.js
Comment on lines +1762 to +1763
const _allowed = new Set();
for (const _t of TIER_ORDER) for (const _m of (_sel.getModelsForTier(_t) || [])) _allowed.add(`${_m.provider}:${_m.model}`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

performance · low
The kNN tier validation loop creates a Set and iterates all tiers every request. If TIER_ORDER has many tiers/models, this could add latency to every routed request.

const msgs = payload?.messages;
if (!Array.isArray(msgs)) return { text: null, index: -1 };
{
const { harnessAskFromPayload } = require('./harness-envelope');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
Same dynamic require pattern here - could cause circular dependency issues if the module loading order changes.

…on, per-decision effort, route preview/audit, corpus gate, evaluation records, cache probe

Declarative routing (config/routing.json, hot-reloaded, validated):
- src/routing/routing-config.js — loader/validator; default config has one
  `legacy` decision (tier from:legacy) so routing is unchanged until edited.
  Harness recognisers (preamble + instruction regex) live here too; the
  Terminus pair is the first entry and harness-envelope reads them.
- src/routing/signals.js — signal registry: anchor_score (banded), jev,
  harness, request_field, tool_count, session_turn, session_phase, risk,
  keyword, legacy_tier, prev_outcome. Pure, fail-open. A condition requires
  its signal to have matched unless matched:false is explicit.
- src/routing/decisions.js — AND/OR/NOT rules, priority, first match wins;
  tier names or from:legacy|anchor|judge; effort/hosts/plugins per decision;
  full trace. observe mode records, enforce mode overrides the legacy tier
  (routing/index.js adopts it and logs a `decision:<name>` escalation).
- config/routing.example.json — Terminus-oriented example: regressing
  sessions escalate, judge/anchor/keyword-hard tasks → COMPLEX@medium pinned
  to DeepSeek, other harness tasks → MEDIUM@low, structured requests keep
  caches/format guard off.

Turn-outcome attribution (src/routing/outcomes.js): on request N+1 classify
turn N as progress / no_progress / regression / provider_error / tool_error /
missing from the evidence the request carries (identical-conversation retry,
repeated command batch, tool_result is_error, environment error markers) plus
the gateway's own record of turn N (telemetry.lastForSession). Only
progress/no_progress/regression are model-attributable; rewardFor() returns
null for environment noise so it never trains the policy. Streaks skip noise.

Per-decision reasoning effort: decisions carry effort; body._effort is
honoured by the OpenRouter and Fireworks invokers ahead of the env maps
(`none` → reasoning disabled on OpenRouter).

Telemetry: decision_name, engine_tier, engine_mode, effort, prev_turn_outcome,
prev_turn_attributable columns; engineFields()/lastForSession() helpers; the
forced-provider hop carries engine + prev_outcome on the body like _jev.

CLI: `lynkr route --preview <req.json> [--json] [--config]` explains every
signal, the legacy pick, shortfall's wish and the engine's decision without
calling a provider; `lynkr audit <session>|--last N` prints the per-turn ledger
(tier, model, decision, effort, latency, cost, judge, previous-turn outcome).
Optional X-Lynkr-Decision* response headers (LYNKR_DECISION_HEADERS=true).

Corpus gate: test/routing-corpus.test.js replays stored signal snapshots
(77 Terminal-Bench first prompts under the example config) through the pure
decision step; any rule/config change that alters a decision fails with the
task and both decisions printed. scripts/build-routing-corpus.js regenerates.

Shortfall: evaluation.records (measured per-model scores) outrank
modelOverrides and seeds; calibrate-capabilities --apply writes them.

scripts/cache-probe.js: two spaced calls per configured host, reports cached
tokens, hit %, cold/warm latency and provider-billed cost (finding: DeepSeek
official 97% hits at 1/12 the cold price; Friendli replica-dependent).

Tests: routing-decisions (8) + routing-corpus (2) added; test:unit 1577 pass,
same 4 pre-existing failures as main; eslint clean on CI scope + new files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment thread bin/lynkr-audit.js
Comment on lines +20 to +21
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');
if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · medium
The audit script reads from a SQLite database file located in the current working directory (.lynkr/telemetry.db) without verifying file permissions or ownership. If the directory has overly permissive permissions, unauthorized users could read sensitive telemetry data including cost, latency, model decisions, and error information.

Suggestion:

Suggested change
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');
if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); }
// Verify database file has appropriate restrictive permissions (e.g., owner-only access)
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');
const dbStats = fs.statSync(dbPath);
if (dbStats.mode & 0o444 || dbStats.mode & 0o222) {
console.error('database file has insecure permissions');
process.exit(2);
}

Comment thread bin/lynkr-audit.js
const lastVal = lastIdx >= 0 ? argv[lastIdx + 1] : null;
const sessionId = argv.find((a) => !a.startsWith('--') && a !== lastVal);
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');
if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
Both CLI scripts call process.exit() without returning from the main() function. While this works, it's inconsistent and could cause issues if the functions are called from tests or other modules.

Suggestion:

Suggested change
if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); }
if (!fs.existsSync(dbPath)) {
console.error(`no telemetry db at ${dbPath}`);
return 2;
}

Comment thread bin/lynkr-audit.js
const rows = db.prepare(`SELECT session_id, COUNT(*) turns, MIN(timestamp) t0, MAX(timestamp) t1, SUM(cost_usd) cost,
SUM(status_code != 200) errors, GROUP_CONCAT(DISTINCT tier) tiers
FROM routing_telemetry WHERE session_id IS NOT NULL GROUP BY session_id ORDER BY t1 DESC LIMIT ?`).all(n);
if (json) { console.log(JSON.stringify(rows, null, 2)); return; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
JSON.stringify is used for terminal output without any post-processing to sanitize sensitive data that could leak into the output object.

Suggestion:

Suggested change
if (json) { console.log(JSON.stringify(rows, null, 2)); return; }
// Sanitize or exclude sensitive fields from output
const sanitizedRows = rows.map(r => {
const { ...rest } = r;
// Remove or redact sensitive fields if needed
return rest;
});
console.log(JSON.stringify(sanitizedRows, null, 2));

Comment thread bin/lynkr-audit.js
const po = r.prev_turn_outcome ? `${r.prev_turn_outcome}${r.prev_turn_attributable ? '' : '(env)'}` : '-';
console.log(` ${pad(i + 1, 3)} ${pad(t, 8)} ${pad(r.tier, 9)} ${pad(String(r.model || '').split('/').pop().slice(0, 30), 30)} ${pad(dec.slice(0, 22), 22)} ${pad(r.effort, 6)} ${pad(r.latency_ms, 7)} ${pad(r.status_code, 4)} ${pad((r.cost_usd || 0).toFixed(5), 9)} ${pad(judge, 14)} ${pad(po, 16)}${r.escalation_source ? ' esc:' + r.escalation_source : ''}`);
});
console.log(`\n cost $${cost.toFixed(4)} in ${inTok} (cached ${cached}, ${inTok ? Math.round(100 * cached / (inTok + cached)) : 0}%) out ${outTok} errors ${errs}\n`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The audit script computes cached percentage using integer arithmetic that could produce confusing results for edge cases (e.g., when cached > input tokens).

Suggestion:

Suggested change
console.log(`\n cost $${cost.toFixed(4)} in ${inTok} (cached ${cached}, ${inTok ? Math.round(100 * cached / (inTok + cached)) : 0}%) out ${outTok} errors ${errs}\n`);
// Add comment explaining the calculation or use clearer logic
console.log(`\n cost $${cost.toFixed(4)} in ${inTok} (cached ${cached}, cached share: ${inTok + cached ? Math.round(100 * cached / (inTok + cached)) : 0}%) out ${outTok} errors ${errs}\n`);

Comment thread bin/lynkr-route.js
console.error('usage: lynkr route --preview <request.json|-> [--json] [--config routing.json]');
process.exit(2);
}
const raw = file === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(file, 'utf8');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · medium
The route script reads request payloads from files or stdin without validating file size limits. A user could provide an extremely large JSON file (e.g., >100MB) that would exhaust available memory when loaded into memory with fs.readFileSync.

Suggestion:

Suggested change
const raw = file === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(file, 'utf8');
// Set a reasonable file size limit (e.g., 10MB) to prevent memory exhaustion
const MAX_FILE_SIZE = 10 * 1024 * 1024;
const stats = fs.statSync(file);
if (stats.size > MAX_FILE_SIZE) {
console.error('input file exceeds maximum size limit (10MB)');
process.exit(2);
}
const raw = fs.readFileSync(file, 'utf8');

Comment on lines +61 to +63
function configPath() {
return process.env.LYNKR_ROUTING_CONFIG ? path.resolve(process.env.LYNKR_ROUTING_CONFIG) : DEFAULT_PATH;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · medium
The LYNKR_ROUTING_CONFIG environment variable is resolved directly via path.resolve() without sanitization. If this env var is user-controllable, it could enable path traversal attacks. Validate the resolved path stays within the expected directory.

Suggestion:

Suggested change
function configPath() {
return process.env.LYNKR_ROUTING_CONFIG ? path.resolve(process.env.LYNKR_ROUTING_CONFIG) : DEFAULT_PATH;
}
function configPath() {
const envPath = process.env.LYNKR_ROUTING_CONFIG;
if (!envPath) return DEFAULT_PATH;
const resolved = path.resolve(envPath);
// Validate resolved path stays within allowed directory
const allowedDir = path.resolve(__dirname, '..', '..', 'config');
if (!resolved.startsWith(allowedDir)) {
logger.warn({ path: resolved }, '[RoutingConfig] invalid config path — using default');
return DEFAULT_PATH;
}
return resolved;
}

Comment thread src/routing/shortfall.js
Comment on lines +315 to +318
for (let i = 0; i < pts.length - 1; i++) {
const [a, va] = pts[i], [b, vb] = pts[i + 1];
if (s >= a && s <= b) {
const f = (s - a) / (b - a);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · medium
The interpolation formula (s - a) / (b - a) risks division by zero if two tier midpoints have the same value. Add a guard to handle this edge case before division.

Suggestion:

Suggested change
for (let i = 0; i < pts.length - 1; i++) {
const [a, va] = pts[i], [b, vb] = pts[i + 1];
if (s >= a && s <= b) {
const f = (s - a) / (b - a);
for (let i = 0; i < pts.length - 1; i++) {
const [a, va] = pts[i], [b, vb] = pts[i + 1];
if (s >= a && s <= b) {
const diff = b - a;
const f = diff === 0 ? 0 : (s - a) / diff;

Comment thread src/routing/shortfall.js
Comment on lines +315 to +318
for (let i = 0; i < pts.length - 1; i++) {
const [a, va] = pts[i], [b, vb] = pts[i + 1];
if (s >= a && s <= b) {
const f = (s - a) / (b - a);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · medium
The interpolation formula (s - a) / (b - a) risks division by zero if two tier midpoints have the same value. Add a guard to handle this edge case before division.

Suggestion:

Suggested change
for (let i = 0; i < pts.length - 1; i++) {
const [a, va] = pts[i], [b, vb] = pts[i + 1];
if (s >= a && s <= b) {
const f = (s - a) / (b - a);
for (let i = 0; i < pts.length - 1; i++) {
const [a, va] = pts[i], [b, vb] = pts[i + 1];
if (s >= a && s <= b) {
const diff = b - a;
const f = diff === 0 ? 0 : (s - a) / diff;

Comment thread src/routing/signals.js
Comment on lines +48 to +49
harness(ctx) {
const { harnessAskFromPayload } = require('./harness-envelope');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The harnessAskFromPayload module is required inside two evaluator functions. While Node.js caches modules, this creates unnecessary require calls and could cause circular dependency issues. Move the require to the top of the file.

Suggestion:

Suggested change
harness(ctx) {
const { harnessAskFromPayload } = require('./harness-envelope');
const { harnessAskFromPayload } = require('./harness-envelope');
const EVALUATORS = {
harness(ctx) {

Comment thread src/routing/signals.js
Comment on lines +48 to +49
harness(ctx) {
const { harnessAskFromPayload } = require('./harness-envelope');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The harnessAskFromPayload module is required inside two evaluator functions. While Node.js caches modules, this creates unnecessary require calls and could cause circular dependency issues. Move the require to the top of the file.

Suggestion:

Suggested change
harness(ctx) {
const { harnessAskFromPayload } = require('./harness-envelope');
const { harnessAskFromPayload } = require('./harness-envelope');
const EVALUATORS = {
harness(ctx) {

veerareddyvishal144 and others added 2 commits October 5, 2026 22:28
…check, cascade trigger

- src/routing/switch-gate.js: hysteresis over attributable turn outcomes
  (outcomes.js ring): escalate after N consecutive regressions/no-progress,
  downgrade after M recoveries above the base tier, min-turns, per-session
  switch budget, cooldown; environment noise never counts. Wraps
  checkSessionPin: observe annotates (gate_action/gate_reason), enforce drops
  the pin and sets a per-session floor the fresh route honours
  (escalation source switch_gate_floor).
- signals.js `complexity`: exemplar contrast — sim(hard centroid) − sim(easy
  centroid) per head with the configured embedder; exemplars from
  config/complexity-exemplars.json (scripts/build-complexity-exemplars.js:
  hard = strong model failed every solo run, easy = cheap model passed).
  scripts/eval-complexity-signal.js reports AUC vs pass/fail for complexity,
  anchor, judge, structural and lifted requirement.
- src/routing/grounding.js: local NLI (Xenova/nli-deberta-v3-small, int8,
  CPU) checks completion claims against the last tool/terminal output;
  verdict supported/unverified/contradicted. Opt-in via LYNKR_GROUNDING
  (observe|flag); runtime loaded lazily from @huggingface/transformers if
  present (not a dependency), fails open. Hooked in invokeModel after the
  provider reply; telemetry grounding_verdict/grounding_contradiction.
  scripts/validate-grounding.js measures it against a run. Measured on v14:
  catches 3/28 false completions, wrongly flags 4/39 — no discriminative
  power on this benchmark (claims describe what the model did and the
  terminal confirms the command ran; hidden tests fail for other reasons).
  Kept observe-only.
- decisions.js cascade trigger: judge top-two margin < trigger_margin marks
  the request ambiguous (telemetry cascade_trigger/cascade_margin); enforce
  path deferred until a verifier proves out.
- telemetry: gate_action, gate_reason, cascade_trigger, cascade_margin,
  grounding_verdict, grounding_contradiction.
- config/routing*.json: switch_gate + cascade (observe), complexity signal;
  example policy uses complexity as a hard-task condition.
- test/switch-gate-grounding.test.js (3 tests). test:unit 1582 pass, same 4
  pre-existing failures; eslint clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…window path

The per-message window loop scores a cleaned single message; the harness
envelope the `harness` signal keys on is stripped there, so the engine
reported `legacy` live while `lynkr route --preview` (full body) matched
the harness decisions. The wrapper now evaluates once on the original body
and forwards that result.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment thread bin/lynkr-audit.js
Comment on lines +21 to +23
if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); }
const Database = require('better-sqlite3');
const db = new Database(dbPath, { readonly: true });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · low
The audit command checks file existence with fs.existsSync() before opening the database, which is good. However, if the database file exists but is empty or has no schema, the Database constructor may still fail. Add more robust error handling around the Database instantiation.

Comment thread bin/lynkr-audit.js
const sessionId = argv.find((a) => !a.startsWith('--') && a !== lastVal);
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');
if (!fs.existsSync(dbPath)) { console.error(`no telemetry db at ${dbPath}`); process.exit(2); }
const Database = require('better-sqlite3');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

test · low
The 'better-sqlite3' package is required at runtime but not verified before use. If this package is not installed in the current environment, the command will crash with a module loading error. Consider adding a try/catch around the require or adding it to package.json dependencies explicitly.

Comment thread bin/lynkr-route.js
// Load the operator .env like the server does; silence logs.
const envPath = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env');
require('dotenv').config({ path: envPath });
process.env.LOG_LEVEL = 'silent'; process.env.LOG_FILE_ENABLED = 'false';

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
Environment variables LOG_LEVEL='silent' and LOG_FILE_ENABLED='false' are set to suppress all logging. While this reduces noise in the preview output, it may also suppress important error diagnostics from the routing module if unexpected issues occur during evaluation.

Comment thread package.json
"lint": "eslint src index.js",
"test": "npm run test:unit && npm run test:performance",
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js",
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js test/shortfall-semantic.test.js test/routing-decisions.test.js test/routing-corpus.test.js test/switch-gate-grounding.test.js",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

test · low
New test files are added to the test:unit script: test/shortfall-semantic.test.js, test/routing-decisions.test.js, test/routing-corpus.test.js, and test/switch-gate-grounding.test.js. These correspond to new routing-related source files introduced in the PR. Ensure these test files exist and are properly implemented before merging.

Suggestion:

Suggested change
"test:unit": "LYNKR_KNN_DIR=/tmp/lynkr-test-knn DATABRICKS_API_KEY=test-key DATABRICKS_API_BASE=http://test.com LOG_FILE_ENABLED=false node --test test/routing.test.js test/hybrid-routing-integration.test.js test/retry-logic.test.js test/sse-transformer.test.js test/passthrough-stream.test.js test/passthrough-mode.test.js test/openrouter-error-resilience.test.js test/format-conversion.test.js test/azure-openai-config.test.js test/azure-openai-format-conversion.test.js test/azure-openai-routing.test.js test/azure-openai-streaming.test.js test/azure-openai-error-resilience.test.js test/azure-openai-integration.test.js test/openai-integration.test.js test/atlas-integration.test.js test/toon-compression.test.js test/gcf-compression.test.js test/llamacpp-integration.test.js test/resilience.test.js test/telemetry-routing.test.js test/memory/store.test.js test/memory/surprise.test.js test/memory/extractor.test.js test/memory/search.test.js test/memory/retriever.test.js test/memory/distiller.test.js test/memory/distiller-freeze.test.js test/memory/wiki.test.js test/memory/skills-cache.test.js test/memory/tencentdb-launcher.test.js test/distill.test.js test/large-payload.test.js test/prompt-cache-injection.test.js test/risk-analyzer.test.js test/interaction-block.test.js test/preflight.test.js test/token-reduction.test.js test/session-affinity.test.js test/cache-state.test.js test/cache-switch-cost.test.js test/lens-recommendations.test.js test/model-registry-cost.test.js test/output-format-guard.test.js test/tier-fallback.test.js test/wrap.test.js test/init.test.js test/tool-call-response-metadata.test.js test/degradation.test.js test/routing-telemetry-columns.test.js test/sticky-routing.test.js test/knn-ambiguous-escalate.test.js test/deescalator.test.js test/client-profiles.test.js test/strip-internal-fields.test.js test/complexity-tool-subtraction.test.js test/bandit.test.js test/routing-propensity.test.js test/reward-pipeline.test.js test/knn-cold-start.test.js test/calibration.test.js test/feedback-loop.test.js test/session-fingerprint.test.js test/side-request-guards.test.js test/verifier.test.js test/intent-score.test.js test/difficulty-classifier.test.js test/classifier-setup.test.js test/usage-stats.test.js test/loop-guard.test.js test/moonshot-model-mapping.test.js test/baidu-model-mapping.test.js test/tenant-policy-ingress-parity.test.js test/decide.test.js test/embeddings-degradation.test.js test/health-probe.test.js test/stuck-detector.test.js test/onnx-embedder.test.js test/ope.test.js test/hierarchical-budget.test.js test/token-rate-limit.test.js test/otel-export.test.js test/mcp-broker.test.js test/compression-budget.test.js test/gpt-utils.test.js test/dedup-observe-only.test.js test/context-window-header.test.js test/token-budget-auto.test.js test/opencode-setup.test.js test/auth-mode-first-party.test.js test/harness-envelope.test.js test/task-ledger.test.js test/jev-router.test.js test/jev-routing.test.js test/force-patterns.test.js test/upstream-fidelity.test.js test/tool-schema-compression.test.js test/passthrough-route.test.js test/quota-ledger.test.js test/shortfall-semantic.test.js test/routing-decisions.test.js test/routing-corpus.test.js test/switch-gate-grounding.test.js",
Verify that the corresponding test files exist in the `test/` directory and contain valid test cases.

Comment on lines +24 to +28
function findRunRoot(dir) {
if (fs.existsSync(path.join(dir, 'results.json'))) return dir;
const subs = fs.readdirSync(dir).filter((d) => fs.existsSync(path.join(dir, d, 'results.json'))).sort().reverse();
return subs.length ? path.join(dir, subs[0]) : null;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The findRunRoot function doesn't verify that directory entries are actually directories before checking for results.json inside them. If a file with a nested path exists, this could throw an error. Add a directory check to be safe.

Suggestion:

Suggested change
function findRunRoot(dir) {
if (fs.existsSync(path.join(dir, 'results.json'))) return dir;
const subs = fs.readdirSync(dir).filter((d) => fs.existsSync(path.join(dir, d, 'results.json'))).sort().reverse();
return subs.length ? path.join(dir, subs[0]) : null;
}
function findRunRoot(dir) {
if (fs.existsSync(path.join(dir, 'results.json'))) return dir;
const subs = fs.readdirSync(dir).filter((d) => fs.statSync(path.join(dir, d)).isDirectory() && fs.existsSync(path.join(dir, d, 'results.json'))).sort().reverse();
return subs.length ? path.join(dir, subs[0]) : null;
}

* Quick check if request should be forced to cloud
*/
function shouldForceCloud(payload) {
if (process.env.FORCE_TIER_PATTERNS === "false") return false;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Hardcoded environment variable 'FORCE_TIER_PATTERNS' is used directly in the logic. Consider centralizing this in routing-config.js or using a shared constants file for environment variable names to maintain consistency across the codebase.

return [{ name: 'harness', preamble: HARNESS_PREAMBLE_RES[0], instruction: HARNESS_INSTRUCTION_RE }];
}

function matchHarness(text) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Harness detection is implemented in multiple places: client-profiles.js uses promptPatterns for pattern detection, while harness-envelope.js uses matchHarness() and isHarnessPrompt(). Consider whether these should share a common implementation or at least use consistent pattern sets.

Comment thread src/routing/index.js
const { buildRequirementVector } = require('./capabilities');
const sf = require('./shortfall');
const req = buildRequirementVector({ dimensions: analysis.breakdown, agenticResult });
const _structuralReq = buildRequirementVector({ dimensions: analysis.breakdown, agenticResult });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Variables _structuralReq and _lifted are declared with underscores suggesting internal use, but they are only used for telemetry logging. If they're needed only for observability, consider using a more descriptive naming convention or documenting why these intermediate values are needed.

const st = fs.statSync(p);
if (_cache.path === p && _cache.mtime === st.mtimeMs) return _cache.config;
const raw = JSON.parse(fs.readFileSync(p, 'utf8'));
const merged = { ...DEFAULT_CONFIG, ...raw, signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) } };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The signal merging on line 120 is shallow: signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) }. This means if you try to override a signal with nested properties (e.g., { anchor: { type: 'anchor_score', extra_option: true } }), you would lose nested properties from DEFAULT_CONFIG. Since the current schema only uses simple { type, ...options } patterns without nested structures, this works today but could break with future schema changes.

Suggestion:

Suggested change
const merged = { ...DEFAULT_CONFIG, ...raw, signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) } };
// Consider a deep merge utility if nested signal options are ever needed
const merged = { ...DEFAULT_CONFIG, ...raw, signals: { ...DEFAULT_CONFIG.signals, ...(raw.signals || {}) } };

Comment on lines +25 to +26
const TTL_MS = 6 * 60 * 60 * 1000;
const _state = new Map(); // sessionId -> { ts, floor, switches, lastSwitchTurn }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
Add a _clear() function call at application shutdown or periodically to drain _state entries that have exceeded max_switches_per_session but haven't hit TTL yet. This prevents memory growth from long-lived but inactive sessions.

Suggestion:

Suggested change
const TTL_MS = 6 * 60 * 60 * 1000;
const _state = new Map(); // sessionId -> { ts, floor, switches, lastSwitchTurn }
const _pinnedProviders = openRouterBody._pinnedProviders; delete openRouterBody._pinnedProviders;
const _schemaOkFn = openRouterBody._schemaOk; delete openRouterBody._schemaOk;
const _timeoutMs = openRouterBody._timeoutMs; delete openRouterBody._timeoutMs;
// ... request code ...
// Restore ALL deleted properties to avoid mutating original body
if (_pinnedProviders) openRouterBody._pinnedProviders = _pinnedProviders;
if (_schemaOkFn) openRouterBody._schemaOk = _schemaOkFn;
if (_timeoutMs) openRouterBody._timeoutMs = _timeoutMs;

…only in enforce mode or when the engine's tier is the served tier

Found during the v15 observe run: effort from a matched decision reached
the invoker even though the engine was in observe mode, so tasks where the
engine and the legacy chain disagreed ran at the engine's effort on the
legacy tier.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment on lines +4 to +7
"signals": {
"anchor": {
"type": "anchor_score"
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
The config/routing.json uses camelCase for signal keys (e.g., tool_count, turn, risk) but config/routing.example.json uses snake_case variants. In src/routing/signals.js, signals are accessed using camelCase keys. This inconsistency in the example config could lead to confusion when using it as a template.

Comment thread config/routing.json
Comment on lines +4 to +8
"anchor_bands": {
"SIMPLE": [
0,
20
],

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

documentation · low
The config/routing.json anchor_bands uses exclusive-high ranges (e.g., SIMPLE: [0, 20] means 0 ≤ score < 20). This is documented implicitly in anchorBandFor() which uses s >= lo && s < hi. Consider adding a comment in the config file to make this contract explicit for future maintainers.

Comment thread config/routing.json
Comment on lines +63 to +68
"complexity": {
"type": "complexity",
"threshold": 0.6,
"scale": 0.15
}
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
The config/routing.json defines a tool_count signal but the complexity signal's bandwidth assignment (line 130 in signals.js) doesn't use it. The signal keys in both config files should be documented to match the signal handlers in src/routing/signals.js.

const requests = [];
for (const entry of fs.readdirSync(from)) {
const p = path.join(from, entry);
if (entry.endsWith('.json') && fs.statSync(p).isFile()) { requests.push({ name: entry.replace(/\.json$/, ''), payload: JSON.parse(fs.readFileSync(p, 'utf8')) }); continue; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

performance · low
build-routing-corpus.js parses JSON files synchronously on line 36. For large directories, consider parallel parsing with Promise.all or streaming JSON parsing.

Comment on lines +112 to +118
function findRunRoot(dir) {
if (!dir || !fs.existsSync(dir)) return null;
if (fs.existsSync(path.join(dir, 'results.json'))) return dir;
const subs = fs.readdirSync(dir).filter((d) => fs.existsSync(path.join(dir, d, 'results.json')))
.sort().reverse();
return subs.length ? path.join(dir, subs[0]) : null;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

other · low
Duplicate logic across multiple scripts: findRunRoot() and expand() are repeated in calibrate-capabilities.js, build-complexity-exemplars.js, build-routing-corpus.js, eval-complexity-signal.js, and validate-grounding.js. Extract these into a shared utility module.

Comment thread src/routing/shortfall.js
if (anchorScore === null || anchorScore === undefined || anchorScore === '') return null;
const s = Number(anchorScore);
if (!Number.isFinite(s)) return null;
const tps = tierProfiles || loadProfiles().tierProfiles;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · high
In anchorRequirement() and jevRequirement(), loadProfiles().tierProfiles is accessed without checking if tierProfiles was explicitly passed as null. If tierProfiles is undefined, loadProfiles() is called which may return null tierProfiles in the error path. Consider adding explicit null checks.

Comment thread src/routing/signals.js
return { matched: n > 0, value: n, confidence: 1 };
},
session_turn(ctx) {
const msgs = Array.isArray(ctx.payload?.messages) ? ctx.payload.messages : [];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · high
In session_phase and complexity evaluators, ctx.payload?.messages is accessed without ensuring ctx.payload exists. Although optional chaining (?.) is used, accessing .filter on an undefined result of ctx.payload?.messages could still fail if the property structure is unexpected. Ensure ctx.payload is checked before accessing messages.

Comment thread src/routing/signals.js
if (!fs.existsSync(file)) return { matched: false, value: null, confidence: null, error: 'no exemplars file' };
const st = fs.statSync(file);
if (!EVALUATORS._cx || EVALUATORS._cx.file !== file || EVALUATORS._cx.mtime !== st.mtimeMs) {
const doc = JSON.parse(fs.readFileSync(file, 'utf8'));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

performance · medium
The complexity evaluator synchronously reads the exemplars file. While cached, this blocks the event loop during initialization. Consider migrating to fs.promises for non-blocking startup.

const TIERS = ['SIMPLE', 'MEDIUM', 'COMPLEX', 'REASONING'];
const PRI = Object.fromEntries(TIERS.map((t, i) => [t, i + 1]));
const TTL_MS = 6 * 60 * 60 * 1000;
const _state = new Map(); // sessionId -> { ts, floor, switches, lastSwitchTurn }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

performance · medium
The global _state Map stores per-session state but relies solely on TTL for cleanup. Long-running processes without active sessions may accumulate stale entries. Consider exposing a periodic cleanup interval or explicit eviction mechanism for production use.

Comment thread src/routing/telemetry.js
prev_turn_attributable: po ? (po.attributable ? 1 : 0) : null,
gate_action: g?.action ?? null,
gate_reason: g?.reason ?? null,
cascade_trigger: e?.cascade ? (e.cascade.triggered ? 1 : 0) : null,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · medium
In engineFields, nested property access like src.engine?.cascade?.triggered uses optional chaining but may still fail if intermediate objects exist but are null. Ensure all levels are properly guarded.

…vLLM Semantic Router

Headline table (same VM, concurrency 8, provider-billed cost), setup for all
three arms, five findings (auto is one model; both routers lose the same
tasks; the gateway decided more than the classifier; Semantic Router lost on
timeouts/failover/per-turn classification; structural difficulty scoring is
flat), what changed in Lynkr, caveats (noise, effort asymmetry, VM change),
and reproduction commands. README gets a one-paragraph pointer next to the
RouterArena result.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Comment thread bin/lynkr-audit.js
const lastIdx = argv.indexOf('--last');
const lastVal = lastIdx >= 0 ? argv[lastIdx + 1] : null;
const sessionId = argv.find((a) => !a.startsWith('--') && a !== lastVal);
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
Hardcoded database path '.lynkr/telemetry.db' may fail in different deployment environments. Consider adding an environment variable override (e.g., LYNKR_TELEMETRY_DB_PATH) or config option similar to how routing.json is handled.

Suggestion:

Suggested change
const dbPath = path.resolve(process.cwd(), '.lynkr', 'telemetry.db');
const dbPath = process.env.LYNKR_TELEMETRY_DB_PATH || path.resolve(process.cwd(), '.lynkr', 'telemetry.db');

Comment thread bin/lynkr-audit.js
const rows = db.prepare(`SELECT session_id, COUNT(*) turns, MIN(timestamp) t0, MAX(timestamp) t1, SUM(cost_usd) cost,
SUM(status_code != 200) errors, GROUP_CONCAT(DISTINCT tier) tiers
FROM routing_telemetry WHERE session_id IS NOT NULL GROUP BY session_id ORDER BY t1 DESC LIMIT ?`).all(n);
if (json) { console.log(JSON.stringify(rows, null, 2)); return; }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · low
Using JSON.stringify on potentially untrusted data in output could expose unexpected data. The tool reads from telemetry database which is controlled, but consider document output if telemetry could contain user-provided fields.

Comment thread bin/lynkr-route.js
}
const raw = file === '-' ? fs.readFileSync(0, 'utf8') : fs.readFileSync(file, 'utf8');
let payload;
try { payload = JSON.parse(raw); } catch (e) { console.error(`not JSON: ${e.message}`); process.exit(2); }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security · medium
JSON.parse() called on user input without structural validation. Consider adding a schema validator to prevent prototype pollution or malformed payloads from causing unexpected behavior.

Suggestion:

Suggested change
try { payload = JSON.parse(raw); } catch (e) { console.error(`not JSON: ${e.message}`); process.exit(2); }
try { payload = JSON.parse(raw); if (!payload || typeof payload !== 'object') { throw new Error('payload must be an object'); } } catch (e) { console.error(`invalid JSON: ${e.message}`); process.exit(2); }

Comment thread bin/lynkr-route.js
}

// Load the operator .env like the server does; silence logs.
const envPath = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Multiple hardcoded .env file paths (cwd then homedir) may be unexpected in containerized environments. Consider adding an environment variable override for consistency with other config paths.

Suggestion:

Suggested change
const envPath = fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env');
const envPath = process.env.LYNKR_DOTENV_PATH || (fs.existsSync(path.join(process.cwd(), '.env')) ? path.join(process.cwd(), '.env') : path.join(require('os').homedir(), '.env'));

Comment on lines +25 to +27
"prev": {
"type": "prev_outcome"
},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · high
The key prev is used in decisions but the actual signal type is prev_outcome (see src/routing/signals.js line 136). This creates a mismatch between the config schema and the signal registry.

Suggestion:

Suggested change
"prev": {
"type": "prev_outcome"
},
"prev_outcome": {
"type": "prev_outcome"
}

Comment thread scripts/cache-probe.js
if (a.error || b.error) { console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ERROR ${a.error || b.error}`); continue; }
const hit = b.prompt ? Math.round(100 * b.cached / b.prompt) : 0;
const note = b.cached === 0 ? 'NO CACHE HITS REPORTED' : (b.ms < a.ms * 0.8 ? 'warm faster' : 'no latency gain');
console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ${String(a.ms).padStart(8)} ${String(b.ms).padStart(8)} ${String(b.cached).padStart(7)} ${String(hit).padStart(5)} ${a.cost != null ? a.cost.toFixed(6).padStart(9) : ' -'} ${b.cost != null ? b.cost.toFixed(6).padStart(9) : ' -'} ${note}`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · low
Inconsistent null comparison: a != null should use strict equality !== null to align with the codebase conventions.

Suggestion:

Suggested change
console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ${String(a.ms).padStart(8)} ${String(b.ms).padStart(8)} ${String(b.cached).padStart(7)} ${String(hit).padStart(5)} ${a.cost != null ? a.cost.toFixed(6).padStart(9) : ' -'} ${b.cost != null ? b.cost.toFixed(6).padStart(9) : ' -'} ${note}`);
console.log(`${model.padEnd(34)} ${String(host).padEnd(12)} ${String(a.ms).padStart(8)} ${String(b.ms).padStart(8)} ${String(b.cached).padStart(7)} ${String(hit).padStart(5)} ${a.cost !== null ? a.cost.toFixed(6).padStart(9) : ' -'} ${b.cost !== null ? b.cost.toFixed(6).padStart(9) : ' -'} ${note}`);

Comment on lines +215 to +216
const r = spawnSync(tb, ['run', '--dataset', opts.dataset, '--agent', 'terminus', '--model', 'anthropic/claude-sonnet-4-5',
'--n-concurrent-trials', String(opts.concurrency), '--output-path', outDir], {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
The lynkr.kill() cleanup only runs if spawnSync completes. If the tb command throws an exception (not a non-zero exit), the orphaned Lynkr process may not be terminated. Consider wrapping the tb spawn in a try-finally to ensure cleanup.

Suggestion:

Suggested change
const r = spawnSync(tb, ['run', '--dataset', opts.dataset, '--agent', 'terminus', '--model', 'anthropic/claude-sonnet-4-5',
'--n-concurrent-trials', String(opts.concurrency), '--output-path', outDir], {
try {
const r = spawnSync(tb, ['run', '--dataset', opts.dataset, '--agent', 'terminus', '--model', 'anthropic/claude-sonnet-4-5',
'--n-concurrent-trials', String(opts.concurrency), '--output-path', outDir], {
env: { ...process.env, ANTHROPIC_API_BASE: `http://localhost:${opts.port}`, ANTHROPIC_API_KEY: 'lynkr-calibration' },
stdio: ['ignore', fs.openSync(path.join(outDir, 'tb-stdout.log'), 'a'), fs.openSync(path.join(outDir, 'tb-stdout.log'), 'a')],
maxBuffer: 1 << 26,
});
// process r.status as before
} finally {
lynkr.kill();
}

function auc(scores, labels) {
// probability that a random FAIL task scores higher (harder) than a random PASS task
const pos = [], neg = [];
scores.forEach((s, i) => { if (s == null) return; (labels[i] ? neg : pos).push(s); });

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
Inconsistent use of null comparison operators. Line 28 uses == null and != null while the codebase otherwise uses strict equality. This is inconsistent with the strict equality requirement in the code standards.

Suggestion:

Suggested change
scores.forEach((s, i) => { if (s == null) return; (labels[i] ? neg : pos).push(s); });
scores.forEach((s, i) => { if (s === null) return; (labels[i] ? neg : pos).push(s); });

for (const p of preds) {
const s = rows.map((r) => r[p]);
const a1 = auc(s, rows.map((r) => r.strongPass)); const a2 = cheap ? auc(s, rows.map((r) => r.cheapPass)) : null;
console.log(`${p.padEnd(12)} ${a1 == null ? ' n/a' : a1.toFixed(3).padStart(15)} ${a2 == null ? ' n/a' : a2.toFixed(3).padStart(10)} ${s.filter((x) => x != null).length}/${rows.length}`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maintainability · medium
Inconsistent use of null comparison operators in logging output.

Suggestion:

Suggested change
console.log(`${p.padEnd(12)} ${a1 == null ? ' n/a' : a1.toFixed(3).padStart(15)} ${a2 == null ? ' n/a' : a2.toFixed(3).padStart(10)} ${s.filter((x) => x != null).length}/${rows.length}`);
console.log(`${p.padEnd(12)} ${a1 === null ? ' n/a' : a1.toFixed(3).padStart(15)} ${a2 === null ? ' n/a' : a2.toFixed(3).padStart(10)} ${s.filter((x) => x !== null).length}/${rows.length}`);

Comment thread src/clients/databricks.js
Comment on lines +944 to +952
for (let ai = 0; ai < _orders.length; ai++) {
const ord = _orders[ai];
if (ord.length) {
openRouterBody.provider = { ...(openRouterBody.provider || {}), order: ord, allow_fallbacks: process.env.OPENROUTER_ALLOW_FALLBACKS === "true" };
const wantSchema = !!(_schemaOkFn && _schemaOkFn(ord[0])) || process.env.OPENROUTER_RESPONSE_FORMAT_SCHEMA === "true";
const rf = openaiResponseFormat(body, wantSchema);
if (rf) { if (rf.type === "json_schema") rf.json_schema.strict = true; openRouterBody.response_format = rf; }
openRouterBody.provider.require_parameters = !!(rf && rf.type === "json_schema");
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bug · medium
The openRouterBody.provider object is modified in place during the fallback loop without restoring previous values between iterations. This means that when trying the next provider, the loop reuses provider settings from the previous attempt. The provider should be reset before each iteration to ensure correct behavior.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant