Skip to content

Close empirical retrieval and generation selection procedure - #30

Draft
oigorbrito wants to merge 40 commits into
rjudi/wave-h-attachment-admission-001from
rjudi/wave-i-empirical-selection-001
Draft

oigorbrito wants to merge 40 commits into
rjudi/wave-h-attachment-admission-001from
rjudi/wave-i-empirical-selection-001

Conversation

@oigorbrito

@oigorbrito oigorbrito commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Wave I — empirical retrieval/generation selection procedure

Base: Wave H head 23c26f7730baec6b05c9231af35406c0d7c69b90.
Exact Wave I head: 6ae960f5cd9ce375a1ec75b025a29cfe51250794.

Scope closed in code:

  • typed retrieval/generation baseline and challenger identities
  • immutable treatment configuration references + SHA-256
  • typed PASS/FAIL/BLOCKED/NOT_TESTED per-case observations
  • typed raw observation artifacts with exact SHA-256 and content binding
  • exact EVAL-010 population binding: 30–50 admitted cases, no subset/extra case, exactly one baseline and one challenger observation per case
  • pre-specified metric ids + direction; undeclared post-hoc metrics rejected
  • non-compensable gates evaluated before numeric comparison
  • strict case×metric Pareto rule
  • NO_CLEAR_WINNER for tradeoffs or equality
  • no weighted score / no post-hoc tie-break
  • manifest SHA, corpus SHA, treatment config SHA, raw observation SHA, dependency SHA and command-evidence SHA verification
  • executed git commit + runtime verification
  • fail-closed string enum encoding
  • external RJ.EmpiricalSelectionVerifier
  • canonical Wave I gate and blocker reconciliation

Treatment taxonomy:

  • R0 = implemented lexical PostgreSQL FTS baseline (ts_rank_cd)
  • R1 = vector-only candidate, NOT_DEMONSTRATED
  • R2 = lexical+vector hybrid candidate, NOT_DEMONSTRATED
  • R3 = hybrid+reranking candidate, NOT_DEMONSTRATED
  • G0 = implemented deterministic fake control only
  • Gx = external generation challenger, EMPIRICAL_DECISION_PENDING

Method classification:

  • paired same-case evaluation and immutable evidence hashes: REPRODUCIBILITY_SUPPORTED
  • pre-specified metrics/hard gates: EMPIRICALLY_SUPPORTED / REPRODUCIBILITY_SUPPORTED
  • prohibition on post-hoc acceptance metrics: DERIVED_FROM_METHOD
  • strict observed-case Pareto rule: PROJECT_DECISION to avoid unsupported weighting
  • R0–R3 and G0/Gx labels: PROJECT_DECISION/spec taxonomy, not empirical findings

Canonical gate:
.\scripts\test-rjudi-wave-i.ps1

External selection evidence required:

  • RJ_EMPIRICAL_SELECTION_MANIFEST_PATH
  • RJ_EMPIRICAL_SELECTION_MANIFEST_SHA256
  • RJ_EMPIRICAL_SELECTION_ARTIFACT_ROOT

Exact-head CI evidence:

  • run 34616947761
  • runner-smoke: failure, steps=null, logs_url=null
  • build-test: skipped
  • no checkout, SDK setup, restore, build, test, PostgreSQL or verifier step executed
  • classification: BLOCKED by RJ-BLK-002, not product FAIL and not PASS

Execution status:

  • local exact-head .NET execution: NOT_TESTED / blocked by RJ-BLK-001
  • remote exact-head code execution: NOT_TESTED / blocked by RJ-BLK-002
  • actual retrieval/generation selection: BLOCKED by RJ-BLK-003

Explicit nonclaims:

  • no vector/hybrid/reranker treatment selected
  • no generation provider/model selected
  • no OpenAI production promotion
  • no weighted aggregate used
  • no cross-case or cross-tribunal quality claim until EVAL-010 is admitted and executed
  • no exact-head PASS until executable evidence exists

Failure rule:

  • pre-step CI provisioning failure = BLOCKED infrastructure, not product FAIL
  • missing EVAL-010/paired evidence = RJ-BLK-003 BLOCKED after executable harness work
  • manifest/hash/raw-binding/runtime/commit mismatch = FAIL until investigated and rerun
  • a valid NO_CLEAR_WINNER is a completed empirical result, not a failure

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant