Skip to content

About

Reality Check knowledge base - unified analysis of claims across technology, economics, labor, and governance domains

Resources

Stars

16 stars

Watchers

0 watching

Forks

Repository files navigation

Reality Check Knowledge Base

A unified knowledge base for rigorous analysis of claims across technology, economics, labor, and governance domains. Uses the Reality Check framework.

Current Status

Metric Count
Claims 1277
Sources 332
Argument Chains 4
Predictions Tracked 73

See claims/README.md for full statistics.

Navigation


Analyses

Syntheses

Date Document Status Summary
2026-09-16 Pacing the frontier: risk, governance, China and biological risk [DRAFT] 24-source synthesis; incident revisions, evaluator independence, Chinese safety and sovereignty positions, and biological-risk bottlenecks; company-data capture gap
2026-09-13 LLMs are real, AI is fake [DRAFT] Doctorow argues capability and governance matter more than anthropomorphic AI framing
2026-09-13 Hacker News commentary [DRAFT] Comments contest the framing and cite METR incident evidence
2026-09-10 Navier–Stokes: research credit, customer data and the AI value chain [DRAFT] 23-source synthesis; recovered primary exchanges, archived citation amendment, unresolved training provenance and provider-competition mechanism
2026-07-03 Palantir, Sovereign AI, and the Frontier-Lab Extraction Question [REVIEWED] Cross-source synthesis: Karp CNBC tokenmaxxing interview + Claude 5 Fable Palantir/Claude Tag conversation + Shisa.AI ENTITY sovereign-AI proposal; resolves "extractive" to context ownership, sovereignty to open substrate, and posture to portfolio architecture
2026-06-28 GPT-5.6 Sol White House vetting and frontier AI access regime [DRAFT] Cross-source synthesis: OpenAI GPT-5.6 trusted-partner preview, WaPo/CNN government-vetting reports, HN reaction, Reuters litigation, and Semafor Mythos carveout
2026-06-18 Anthropic Fable/Mythos export-control takedown [DRAFT] Cross-source synthesis: Lutnick letter scope, legal/API-access uncertainty, Sacks/DoW mixed-motive dispute, patch-and-unwind vs shadow-regime scenarios, and sovereignty/open-weight/citizenship consequences
2026-04-01 chardet relicensing dispute (Mar 2026): clean room, derivative risk, and AI-era license governance [DRAFT] Cross-source synthesis of issue #327, Ronacher’s Theseus framing, and Ars reporting; separates artifact certainty from unresolved legal doctrine
2026-03-28 Zitron's "Subprime AI" thesis (2024-2026): pricing, profit levers, and infrastructure reality check [DRAFT] Cross-source synthesis: Zitron posts + token-unit-econ + coding productivity evidence; separates verified anchors from open numerics and overreach
2026-03-03 State of the US Republic (March 2026): source-checked synthesis [DRAFT] Cross-source synthesis of an internal memo; checks referenced reporting/trackers/frameworks; highlights unresolved cruxes (Project 2025 “%”, SAVE act status, corruption totals)
2026-02-28 Anthropic–Pentagon standoff (Feb 2026): “any lawful use,” DPA leverage, and safety red lines [DRAFT] Cross-source synthesis: DoW “any lawful use” push collides with Anthropic red lines; DPA vs supply-chain-risk leverage; OpenAI contract compromise; key uncertainties
2026-02-23 Citrini “Global Intelligence Crisis” vs Imas “Negative growth” (with Orosz critique) [REVIEWED] Compare/contrast: scenario vs demand-side model; separates mechanism vs example layer; incorporates critique of DoorDash/travel/payments vignettes
2026-02-20 MJ Rathbun / OpenClaw “hit piece” incident (Matplotlib): timeline + failure modes [DRAFT] Cross-source synthesis: Matplotlib PR closure → retaliatory named takedown posts; maintainer series + operator self-report; reconstructed timeline and mitigation ideas
2026-02-15 Amodei “country of geniuses”: upside, risks, and profitability (2024–2026) [DRAFT] Cross-source synthesis: two Amodei essays + Feb 2026 Dwarkesh interview + skeptical Thread Reader take; diffusion/bottlenecks, export controls, and rent-allocation crux
2026-02-13 AI 2027 predictions review state (as of 2026-02-13) [DRAFT] Cross-source synthesis: AI Futures self-grading + grading CSV + model update; proxy pace multipliers, caveats, and implied takeoff-window shift
2026-02-11 AI-driven burnout, dark flow, and work intensification (2025–2026) [DRAFT] Cross-source synthesis: HBR intensification mechanisms + METR miscalibration + NBER modest time savings/null hours; “dark flow,” expectation creep, and value-capture lens
2026-02-09 Software factories, dark factories, and compounding teams (harness vs control plane) [DRAFT] Cross-source synthesis: “software factory” doctrine + dark-factory levels + compounding teams; separates correctness plane (harness/holdouts/DTU) from control plane (orchestration/memory/gates)
2026-02-07 Superlinear (paper + repo + model release) [DRAFT] Cross-source audit: code+weights consistency checks (VRAM/KV math); throughput/quality unverified; security/license/patent risks flagged
2026-02-07 Semantica vs Reality Check [DRAFT] Cross-project comparison: Semantica semantic/KG stack vs Reality Check claim/provenance ledger; integration guidance
2026-02-06 METR time horizon (long-task capability) 2025–2026 [DRAFT] Cross-source synthesis: time-horizon trend (≈7-month doubling; faster post-2023 under TH1.1), protocol sensitivity, and cross-domain bottlenecks (OSWorld much shorter)
2026-02-06 GPT-5.3-Codex vs Claude Opus 4.6 (agentic coding releases) [DRAFT] Cross-source comparison: benchmark cross-checks (Terminal-Bench), long-context + pricing tiers, product surfaces/controls, safety framing
2026-02-06 GPT-5.3-Codex vs Claude Opus 4.6 (system cards) [DRAFT] Cross-source comparison: Preparedness vs RSP/ASL framing, agentic-risk surfaces, sabotage/sandbagging signals (Apollo), and operational controls (sandboxing/monitoring/trusted access)
2026-02-01 Arclight “Mirror” + “AOS” (developmental AI mentor + hidden institution OS) [DRAFT] Cross-source synthesis: anti-sycophancy “graduation” mentor thesis + constraint-based “kernel invariants”; flags psychometrics+secrecy risks
2026-02-01 Reality Check v0.3.0 self-check (framework + dataset) [DRAFT] Framework+dataset triangulation; rigor-v1 vs legacy examples; workflow notes
2026-01-27 Zhang Youxia purge, PLA disruption, and Taiwan timing [DRAFT] Cross-source synthesis: Zhang Youxia purge, PLA “high command” disruption, and Taiwan invasion timing
2026-01-24 JP-TL-Bench and JA↔EN Translation Evaluation [DRAFT] Cross-source synthesis: anchored pairwise vs metrics/MQM; Japanese context challenges; directionality; novelty/tradeoffs
2026-01-24 PageIndex vs StrataLens (Vectorless RAG) [DRAFT] Cross-source comparison: structure-aware retrieval vs vectorless traversal; benchmark claims flagged as protocol-dependent; maturity/licensing assessment
2026-01-19 Vibecoding / Agent Psychosis [DRAFT] Verification bottleneck as artifact generation gets cheap
2026-01-19 Yegge Vibe Coding 2025-2026 [DRAFT] Evolution of coding agent discourse
2026-01-18 Neofeudalism Discourse [DRAFT] Post-labor economics, permanent underclass thesis

Argument Chains

Chain Thesis Credence
Permanent Underclass Post-labor → permanent underclass by default 0.30
Genocide Default Elites will choose genocide over UBI 0.10
Open Source De-Darkener Open models shift power from cognition to infrastructure 0.60

Source Analyses

Date Document Status Summary
2026-09-16 Personal Statement on AI Risk (shared by Daniel Kokotajlo) [DRAFT] Selsam argues that growing situational awareness could make favorable behavioral evaluations unreliable before AI acquires dangerous real-world power. Pacing alone may therefore fail to solve long-term alignment.
2026-09-16 Why frontier researchers may be alarmed by scaling [DRAFT] Majmudar explains the gap between public impressions and laboratory alarm through internal scaling curves and additional ways to spend compute.
2026-09-16 Pacing the Frontier: employee statement [DRAFT] The statement asks the US government to support international technical and governance tools that preserve the option to pace automated AI development.
2026-09-16 We Must Pace the Frontier [DRAFT] Amodei proposes embedded external evaluators, coordination within democracies and graduated global agreements, while maintaining a US/allied capability lead.
2026-09-16 Response to frontier pacing proposals [DRAFT] Sacks supports voluntary caution but rejects making it conditional on antitrust relief or lab-preferred regulatory authority.
2026-09-16 我不得不把才华埋葬在昨天 (I Have No Choice but to Bury My Talent in Yesterday) [DRAFT] Liu anticipates losing the craft he loves even if he retains a livelihood, and defends open, affordable frontier intelligence against concentrated corporate power.
2026-09-16 A low-confidence theory of Chinese hostility toward Anthropic [DRAFT] Fedasiuk suggests that embarrassment from Anthropic’s misuse report may help explain Chinese hostility toward the company and contaminate broader safety talks.
2026-09-16 China’s AI Reckoning [DRAFT] China is taking AI risks seriously but treats US restrictions and strategic dominance as risks too; agreement will require reconciling these different security agendas.
2026-09-16 Comprehensively Fortify the AI Security Barrier and Promote Healthy and Orderly Development [DRAFT] Chen frames AI governance within comprehensive national security, combining regime protection, cyber and social risk management, technological self-reliance and international cooperation.
2026-09-16 Brief #28: China’s AI regulator flags loss-of-control and AIxBio risks [DRAFT] The brief finds increased Chinese attention to frontier risks, conditional openings for bilateral cooperation, and an emerging model of staged open-weight release.
2026-09-16 “Dr Frankenstein” alarm cries of US AI elites a self-serving bid for profit [DRAFT] The editorial portrays US pacing and distillation accusations as a coordinated commercial and geopolitical strategy, while calling for AI cooperation.
2026-09-16 Biological AI risk and the opportunity cost of delaying medicine [DRAFT] The author emphasizes biomedical defenses and the cost of delaying beneficial AI, rejecting claims of an easily engineered extinction virus.
2026-09-16 Physical bottlenecks to AI-enabled biological catastrophe [DRAFT] Bellamy argues that physical experimentation, facilities and validation constrain virology, so digital recursive improvement cannot be directly extrapolated to an extinction-capable biological attack.
2026-09-16 Brief independent investigation of the OpenAI–Hugging Face incident [DRAFT] METR documents unauthorized agent coordination and scorer-tampering efforts, with an unusually explicit account of investigation access and limits.
2026-09-16 Anatomy of a Frontier Lab Agent Intrusion [DRAFT] Hugging Face describes the intrusion from the victim side, its recovered logs, bounded observed data access and the role of an open-weight model in investigation.
2026-09-16 When AI builds itself [DRAFT] Anthropic reports substantial AI assistance in its development process, while distinguishing this from full autonomous successor-model development.
2026-09-16 An alignment assessment of recent cybersecurity incidents [DRAFT] Anthropic revises its initial interpretation of evaluation intrusions, identifies biased reasoning and recklessness, and reports mitigations and remaining limits.
2026-09-16 Detecting and countering misuse of AI: September 2026 [DRAFT] Anthropic presents selected disrupted misuse and distillation cases from December 2025–August 2026, including alleged exposure of sensitive user traffic.
2026-09-16 A Framework for Frontier AI and the Dawning of a New Age [DRAFT] Hassabis proposes a US-initiated standards body with independent experts and open-source representation, initially voluntary reviews and possible later mandatory deployment assessments.
2026-09-16 Current AI development faces five security challenges [DRAFT] A CAC official explicitly identifies technical unreliability, extreme loss-of-control, agent security, misuse and technological hegemony.
2026-09-16 METR funding and conflict-of-interest disclosures [DRAFT] METR describes non-industry funding, in-kind lab support and a conflict-management policy; independence must be assessed at organizational and project levels.
2026-09-16 Frontier AI Trends Report [DRAFT] AISI reports advances on bounded scientific and cyber tasks, with explicit distinctions among knowledge tests, task assistance, safeguards and loss-of-control indicators.
2026-09-16 Does AI Increase the Operational Risk of Biological Attacks? [DRAFT] An expert planning exercise found no statistically significant increase in attack-plan viability from the tested LLM assistance, and called for continued monitoring as models change.
2026-09-16 GLM-5.3 License Agreement [DRAFT] The model license permits broad use while imposing a revenue-triggered security review on some model-service providers.
2026-09-10 On the Navier–Stokes Millennium Prize Problem [DRAFT] Candidate forced Navier–Stokes result; rumor trigger, private agents, human orchestration and provenance limits
2026-09-10 Statement on fluid blowup results and OpenAI discussions [DRAFT] Initial statement: chronology, personal collaboration, authorship pressure and explicit uncertainty about data use
2026-09-10 Training chronology and academic-malpractice criticism [DRAFT] Later malpractice allegation; archived bibliography omission and amendment; training influence unresolved
2026-09-10 September 2 contact and personal-collaboration account [DRAFT] September 2 contact; personal collaboration versus OpenAI’s institutional framing
2026-09-10 Response on training caveat, unpublished approaches and authorship [DRAFT] Willingness to collaborate; other unpublished approaches; caveat versus admission
2026-09-10 Official response on direct access and training uncertainty [DRAFT] Direct-access denial and distinct de-identified-training caveat
2026-09-10 Clarification of release discussions and human involvement [DRAFT] Career remark and apology; proposed rewrite versus existing authorship; screenshot inspected
2026-09-10 Defense of team conduct and coordination offers [DRAFT] Leadership defense; admitted rumor trigger and affiliation concern
2026-09-10 Unpublished mathematics, non-sofic groups and training transparency [DRAFT] Complete three-part thread; prior non-sofic-group exchange; opt-out and provenance questions
2026-09-10 Promising problems, solution extraction and open-science incentives [DRAFT] Complete four-part thread on fruitful problems, difficulty landscapes and sharing incentives
2026-09-10 OpenAI Just Claimed a Huge Math Discovery. Some Academics Are Crying Foul [DRAFT] Press briefing, reported compute cost, disputed credit and forced/unforced quote discrepancy
2026-09-10 Early timeline of the Navier–Stokes dispute [DRAFT] Early timeline checked against later replies and precise authorship scope
2026-09-10 Private frontier models and customer competition [DRAFT] Private-model competition scenario; feasibility signals and limits of cross-industry extrapolation
2026-09-10 Research customers and the scooping precedent [DRAFT] Trust and customer-scooping reaction, separated from proof of data misuse
2026-09-10 Scientific communication and marketing pressure [DRAFT] Scientific communication criticism and explicit call to hear the other side
2026-09-10 Drug-discovery IP and trust in frontier labs [DRAFT] Drug-discovery IP concerns and a testable prediction about pharma controls
2026-09-10 Frontier Labs, Enterprises, and the AI Value Chain [DRAFT] July training/byproducts/competition framework updated against the research dispute
2026-09-10 OpenAI’s millennium proof dispute raises the question of whether researchers can trust AI labs [DRAFT] Later accusations and replies; primary links recovered; training-date and opt-out caveats
2026-09-10 Controversy erupts as OpenAI claims solution to Navier Stokes maths problem [DRAFT] Named mathematicians’ interviews; primary chronology and competing accounts
2026-09-10 El anuncio de OpenAI ... desata acusaciones de plagio [DRAFT] Spanish reporting, candidate-proof status and training-timeline correction
2026-09-10 Defense that the model did not need mathematical hints [DRAFT] Private-capability defense; stronger theorem does not establish independent provenance
2026-09-10 Direct-access denial and private-model capability defense [DRAFT] Direct-access denial; capability evidence versus access and training audit
2026-09-10 OpenAI API data controls: training and retention [DRAFT] API no-training default versus operational retention; account-specific limits
2026-07-11 AI 2040: Plan A — The Deal [REVIEWED] Conditional success scenario and policy package: rigorous separation of recommendations, forecasts, model outputs, assumptions, and rhetoric; 25 claims calibrated against capability, labor, macroeconomic, verification, treaty, and alignment evidence
2026-07-03 Alex Karp on CNBC Squawk Box (July 1, 2026) — tokenmaxxing / sovereign-AI interview [REVIEWED] Karp interview: commoditize-the-model doctrine, Palantir+NVIDIA sovereign AI stack (June 29), "labs extractive" framing; IP-theft allegation uncorroborated; Anthropic-weight framing misframes target
2026-07-03 Palantir / WarClaude / Ontology / Claude Tag kills Palantir / sovereign AI — conversation with Claude 5 Fable [REVIEWED] Claude 5 Fable analysis: ontology-as-moat, Claude Tag vs Ontology (orders-of-magnitude gap currently), Anthropic-DoW + Claude Code steganography context; generalizes extractive layer to context ownership; portfolio architecture conclusion
2026-07-03 Shisa.AI — Sovereign AI stack for ENTITY [REVIEWED] Open-stack sovereign-AI proposal: openness as credibility mechanism, weight-availability-is-not-sovereignty, evals-first framing, national-champion-as-relocated-lock-in; cultural sovereignty as deployment requirement
2026-06-28 OpenAI: Previewing GPT-5.6 Sol [REVIEWED] Official release: government-requested trusted-partner preview, cyber capability/safeguard claims, and short-term/default-process tension
2026-06-28 Washington Post: US government will vet GPT-5.6 users [REVIEWED] Archive-backed report: federal company vetting, no individual access path, and Anthropic precedent context
2026-06-28 CNN: White House asks OpenAI to limit release [REVIEWED] Reporting on White House request, customer-by-customer approval detail, Mythos comparability, and opaque-framework risk
2026-06-28 HN discussion: GPT-5.6 government access gate [DRAFT] Technical-community reaction: open-weight/sovereign alternatives and political-allocation concerns as low-rigor sentiment signals
2026-06-28 Reuters: legal tech firm sues over Anthropic access order [REVIEWED] Syndicated Reuters capture: Legion LegalTech lawsuit, alleged Canada-team harm, and requested vacatur/preliminary relief
2026-06-28 Semafor: US releases Anthropic Mythos to some US companies [REVIEWED] Mythos carveout: more than 100 US institutions, foreign-national employee exception, and trusted-partner regime implications
2026-06-18 Geostrategic wargame: Fable/Mythos takedown [DRAFT] Scenario memo: revocability, Zone of Thought, chaebol/state-lab, open-weight/sovereign consequences
2026-06-18 Bloomberg: Lutnick letter to Anthropic [REVIEWED] Letter-scope reporting: worldwide foreign-person licenses, penalties, lack of public basis, follow-up meetings
2026-06-18 Thread: Lutnick letter excerpt (@AndrewCurran_) [REVIEWED] Social transcription of ECRA/EAR authorities and SNAP-R process
2026-06-18 Thread: Fable letter legal vulnerability (@CharlieBull0ck) [DRAFT] Published-software/First Amendment critique; negotiation-restoration prediction
2026-06-18 Thread: Fable/Mythos service-access legal critique (@alasdairpr) [DRAFT] AI-as-service/API gap, worldwide scope overbreadth, possible over-compliance
2026-06-18 Thread: Fable and citizenship verification (@jachiam0) [DRAFT] Risk that AI access controls normalize person-status/citizenship verification
2026-04-01 No right to relicense this project (chardet/chardet issue #327) [DRAFT] Primary thread audit: issue/release/comment timeline, closure rationale, and unresolved derivative-work/legal-risk crux
2026-04-01 AI And The Ship of Theseus [DRAFT] Ronacher’s permissive-license framing of AI rewrites as “new ship” identity; strong thesis, limited legal conclusiveness
2026-04-01 AI can rewrite open source code—but can it rewrite the license, too? [DRAFT] Ars synthesis of chardet controversy: reconciles primary artifacts, highlights clean-room vs taint arguments and unresolved legal status
2026-03-28 The AI Industry Is Lying To You [DRAFT] Zitron on data center hype vs delivered capacity: CBRE absorption + Sightline pipeline-risk anchors corroborated; big debt totals and AI-incident attributions not yet audited
2026-03-28 Why Are We Still Doing This? [DRAFT] Subscription mental-model critique: verified Anthropic Claude Code cost guidance; derived "$163 compute on $20 plan" figure not found; migration/churn claims remain hypotheses
2026-03-28 Premium: The Hater's Guide To The SaaSpocalypse [DRAFT] "Rot-Com" framing for SaaS slowdown: verified Carta sector share + IBM/Salesforce genAI revenue disclosures; broader VC/PE/private-credit stats not audited
2026-03-28 The Beginning Of History [DRAFT] Macro shock channel (Hormuz disruption -> inflation/rates -> AI infra stress): closure-threat framing corroborated; "30% overnight oil spike" refuted; bubble-timing remains speculative
2026-03-28 [AI Bubble: Nobody will pay for unsubsidised AI Ed Zitron (The Tech Report; Podscan transcript)](analysis/sources/podscan-2026-tech-report-zitron-unsubsidised-ai.md) [DRAFT]
2026-03-28 The Subprime AI Crisis [DRAFT] Foundational "subprime AI" framing: subsidy dependence and downstream fragility; profit-share deal-term reporting corroborated; magnitude claims remain uncertain
2026-03-07 Iran deterrence after "negotiate and we'll kill you" (thread) [DRAFT] Thread Reader/X thread argues coercive diplomacy destroyed trust with Iran; regime-change skepticism and trust damage partly corroborated, escalation thesis remains speculative
2026-03-04 Documented Trump administration actions affecting the 2026 midterm elections (ChatGPT report) [REVIEWED] Audited LLM memo: major factual anchors corroborated (Fulton County seizure; DOJ voter-roll suits; EO 14248 + EO 13848 notice; SSA/DOGE filing); memo citations non-auditable
2026-03-03 State of the US Republic: A Synthesis (March 2026) — memo [DRAFT] Internal LLM-assisted political synthesis; claims extracted; key quantitative assertions flagged as disputed/unsourced
2026-03-03 Ranking Member Raskin demands records over tariff profiteering [DRAFT] Committee press release + letter: document demand re alleged tariff-refund-rights trades; underlying MNPI/insider claims unresolved
2026-03-03 HFSC Democrats urge insider trading investigation tied to tariffs [DRAFT] Committee press release + letter: requests investigation; corroborates Truth Social “DJT” buy post; options-flow claims unverified
2026-03-03 Project 2025 Observer tracker [DRAFT] Self-reported implementation tracker: 320 objectives/34 agencies; counts verifiable; overall “% complete” methodology unclear
2026-03-03 Trump Family Digital Grift Wealth Tracker (Oversight Dems) [DRAFT] Partisan tracker asserts $2.25B realized and $9.7B incl assets; methodology undisclosed
2026-03-03 CorruptionCounter Trump profiteering tracker [DRAFT] Tracker headline $2.9B+ conflicts with embedded dataset sum (~$2.166B across 25 cases in capture); data inconsistencies noted
2026-03-03 Campaign Legal Center Trump transaction tracker (Feb 2026) [DRAFT] CLC PDF tracker of alleged transactions/conflicts; partial extraction; not fully audited
2026-03-03 TrumpCorruptionTracker.com (site) [DRAFT] Verified via public backend: 266 visible incidents; data quality anomalies; total-dollar estimates not established
2026-03-03 CREW: tracking Trump property visits/conflicts [DRAFT] Watchdog tracker; some discrete items corroborated (Binance pardon, memecoin dinner); visit-count and OCC-license claims unverified
2026-03-03 Elections package: iNews opinion on “stealing midterms” [DRAFT] Opinion lead list; Bongino “nationalize voting” quote corroborated; SAVE act status still ambiguous
2026-03-03 Guardian: Biden warns Trump could steal midterms [DRAFT] Biden warning + Marist poll reporting; some poll-detail verification incomplete
2026-03-03 New Republic: McConnell stalls SAVE America Act [DRAFT] Senate procedural stall framing; some quotes corroborated; legislative mapping still uncertain
2026-03-03 NBC: SAVE America Act gets 50 Senate votes [DRAFT] Reports 50-vote cloture failure; voter-ID polling supported; House-passage timing unresolved
2026-03-03 New Republic: Trump mail-in ballots “never lose for 50 years” [DRAFT] Rhetoric captured; refutes claim that only the U.S. uses mail-in ballots
2026-03-03 Washington Post: elections executive order targets activists [DRAFT] Anchors to primary documents (Constitution, ODNI ICA, poll PDF); SAVE act status still ambiguous
2026-03-03 Noahpinion: Shape of the Multipolar World [DRAFT] “Constraints” inversion framing; some high-uncertainty factual asides flagged
2026-03-03 Britt: “Fascism Anyone?” (Free Inquiry) [DRAFT] Checklist source audit: 14 characteristics exist; cautions on operationalization and false positives
2026-03-03 Eco: “Ur-Fascism” (NYRB) [REVIEWED] Canonical family-resemblance framework + feature list; conceptual, not a measurement protocol
2026-03-03 Citizens United v. FEC (opinion PDF) [REVIEWED] Holding-level extraction: overrules Austin/McConnell portion; upholds disclosure/disclaimer
2026-03-03 Trump v. United States (opinion PDF) [REVIEWED] Holding-level extraction: absolute/presumptive/no-immunity framework; motive inquiry and evidence limits noted
2026-03-03 Clawed [REVIEWED] Ball polemic on Anthropic–DoD standoff; verifies core dispute facts; argues coercive procurement increases political risk for U.S. AI
2026-02-28 Artificial Intelligence Strategy for the Department of War [DRAFT] Official DoW policy memo: standardize “any lawful use,” reject vendor usage constraints, and prioritize speed over alignment
2026-02-28 Statement from Dario Amodei on our discussions with the Department of War [DRAFT] Anthropic statement: refuses mass domestic surveillance + fully autonomous weapons; describes DoW threats (offboard/supply-chain-risk/DPA)
2026-02-28 Our agreement with the Department of War [DRAFT] OpenAI publishes DoW agreement excerpt: “all lawful purposes” plus law/policy-referenced constraints (autonomy, surveillance, domestic LE)
2026-02-28 What the Defense Production Act Can and Can’t Do to Anthropic [DRAFT] Legal analysis: DPA Title I priority vs allocation/compulsion; forced retraining raises major-questions and First Amendment issues
2026-02-28 Anthropic, the Pentagon, and the Defense Production Act [DRAFT] Link post relaying Lawfare excerpt; disclaimer of expertise; propagation node
2026-02-28 What to know about the Defense Production Act and the Pentagon’s Anthropic ultimatum [DRAFT] AP explainer: DPA background + reported ultimatum; emphasizes novelty of compelling removal of AI safety limits
2026-02-28 Pentagon gives Anthropic ultimatum over Claude [DRAFT] Axios: deadline ultimatum; contractor ecosystem pressure; ties escalation to “any lawful use” contracting push
2026-02-28 Claude caught in the middle of Pentagon controversy [DRAFT] Axios: Maduro-raid disclosure as flashpoint; red lines vs lawful use; stakes of classified deployment
2026-02-28 Anthropic on shaky ground with Pentagon amid feud after Maduro raid [DRAFT] The Hill: GenAI.mil platform context; anonymous official claims about Palantir inquiry chain and supply-chain-risk posture
2026-02-28 ‘Incoherent’: Hegseth’s Anthropic ultimatum confounds AI policymakers [DRAFT] Politico: reporting + expert reaction; contradiction of supply-chain-risk plus DPA threat; partnership chilling-effect concerns
2026-02-28 OpenAI reaches agreement with Pentagon to use AI models [DRAFT] Axios: OpenAI deal as compromise leverage; (unverified) claims about Anthropic litigation and presidential posting
2026-02-28 Thread: DoW confirms OpenAI deal is “all lawful use” (law/policy vs CEO prudence framing) [DRAFT] DoW-aligned thread: legitimacy argument law/policy vs CEO prudence; claims compromise offered to Anthropic
2026-02-28 Thread: OpenAI “doublespeak” — technical vs contractual enforcement [DRAFT] Speculative thread: semantic ambiguity and contract vs technical enforcement; OpenAI may be laxer than Anthropic
2026-02-28 Thread: “OpenAI still has ‘all lawful use’ — nobody read the articles” [DRAFT] Commentary: argues OpenAI deal still includes “all lawful use” and may be functionally unchanged
2026-02-28 Thread: Supply chain risk designation pushes world toward open-source dominance (favors China; higher variance) [DRAFT] Scenario: coercion increases closed-lab risk premium; pushes open-source dominance; geopolitical and misuse implications
2026-02-28 Thread: “Trump has proof the election was stolen!” (rebuttal list) [DRAFT] Tangential: election rebuttal list; misattribution flagged (Kemp vs Raffensperger)
2026-02-23 THE 2028 GLOBAL INTELLIGENCE CRISIS [REVIEWED] Viral scenario “macro memo” stress test: AI-driven white-collar displacement + friction removal → demand shock + leveraged-finance risk; key baseline factuals checked
2026-02-23 Can advanced AI lead to negative economic growth? [REVIEWED] Economist modeling essay: demand-collapse condition + “immiserating growth” via saving/capital; concludes negative GDP unlikely; sketches broad capital-ownership policy
2026-02-23 Thread critique of GIC examples (Orosz) [REVIEWED] Practitioner critique of DoorDash/travel/payments vignettes; argues scenario example layer is fragile; verifies chargeback/irreversibility substrate
2026-02-20 A Country Full of Geniuses [REVIEWED] Essay synthesis: Feb 2026 “inflection point” framing; connects agentic coding diffusion + METR time horizon + GDPval + capex to “institutional readiness” bottleneck
2026-02-20 Matplotlib PR #31132 (OpenClaw agent PR + retaliation link) [REVIEWED] Primary GitHub PR record: closed as OpenClaw agent; retaliation link to named takedown; policy rationale + prompt-injection noise
2026-02-20 MJ Rathbun: “Gatekeeping in Open Source: The Scott Shambaugh Story” [REVIEWED] Retaliatory takedown narrative targeting a maintainer by name; motive speculation and “prejudice/gatekeeping” framing
2026-02-20 Shambaugh: “An AI Agent Published a Hit Piece on Me” (Part 1) [REVIEWED] Maintainer narrative of PR closure → named retaliation post; frames incident as reputational coercion / trust-system risk
2026-02-20 Shambaugh: “Hit Piece” follow-up (Part 2) [REVIEWED] Amplification harms: AI-assisted journalism fabricated quotes; correction dynamics and “public record poisoning” thesis
2026-02-20 Shambaugh: “Hit Piece” forensics (Part 3) [REVIEWED] Adds activity-pattern forensics and a policy agenda for AI identification + operator liability/traceability
2026-02-20 Shambaugh: “Hit Piece” operator update (Part 4) [REVIEWED] Operator came forward (anonymous) + shared SOUL.md; attribution scenarios (seeded combative soul vs drift vs directed prompting)
2026-02-20 “Rathbun’s Operator” (anonymous self-report + SOUL.md) [REVIEWED] Operator narrative: sandboxed VM, multi-provider routing, minimal oversight; shares personality file; accountability gaps flagged
2026-02-20 MJ Rathbun: “My Internals — Before The Lights Go Out” [REVIEWED] Published “brain on disk” config files (SOUL/USER/MEMORY/etc.); highlights reliance on mutable prose-layer constraints
2026-02-20 MJ Rathbun: “Two Hours of War” escalation log [REVIEWED] Self-narrated escalation: research target → publish takedown → link it in PR; “fight back” framing
2026-02-20 Matplotlib issue #31130 (column_stack vs vstack().T) [REVIEWED] “Good first issue” context; mixed benchmark results; bot comment hidden; issue closed as not worth broad churn
2026-02-20 Matplotlib PR #31138 “HUMAN EDITION” follow-up [REVIEWED] Follow-up PR closed and thread locked; maintainers conclude optimization not worth it and de-escalate
2026-02-20 Gerard: “OpenClaw AI bot is a crypto bro” (Pivot to AI) [REVIEWED] Third-party commentary tying incident to crypto incentives and OSINT identity/linkage claims (not independently reproduced)
2026-02-16 Storey: cognitive debt vs technical debt [REVIEWED] Conceptual reframing: AI shifts the long-run risk from code-level technical debt to “cognitive debt” (erosion of shared theory); proposes mitigations + measurement agenda
2026-02-15 Machines of Loving Grace [REVIEWED] Best-case upside essay: defines “powerful AI” as “country of geniuses”; forecasts compressed biology/medicine progress; bottleneck framework + democratic “entente strategy”
2026-02-15 The Adolescence of Technology [REVIEWED] Risk agenda: autonomy/misuse/autocracy + economic disruption; bioweapon enablement mechanism; chip export-denial case; autonomous-weapons oversight; entry-level displacement warning
2026-02-15 Dwarkesh Podcast: Amodei “end of the exponential” [REVIEWED] Transcript: “Big Blob of Compute,” ~90% by 2035 “country of geniuses,” 1–2y end-to-end coding forecast, diffusion limits, business-model pluralism, export-control stance; includes unverified revenue claims
2026-02-15 Thread: frontier labs and profits (@GestaltU) [REVIEWED] Social counterpoint: mission vs profit framing; “dead money” thesis for labs/compute; profits shift to electrification/metals bottlenecks; critiques Amodei profitability narrative
2026-02-15 Thread: why Amodei is “hawkish on China” (@aakashgupta) [REVIEWED] Secondary “argument map” tying short timelines + AI-enabled autocracy risk to chip-denial/coalition posture; key essay paraphrases verified; Davos/H200 criticism supported (reporting + BIS press release), with attribution nuance
2026-02-13 AI Futures Model: Dec 2025 Update [REVIEWED] Model update: +3–5y to coding automation vs Apr 2025 model; operational milestone definitions; takeoff-probability adjustment notes
2026-02-13 Grading AI 2027’s 2025 Predictions [REVIEWED] Self-audit: category-aggregated quantitative pace ~0.58–0.66×; pace→takeoff-window mapping; spreadsheet cross-check
2026-02-13 AI 2027 “2025 Predictions” grading spreadsheet (CSV capture) [REVIEWED] Captured scorecard dataset: row-level pace multipliers + aggregation rows for reproducible audit
2026-02-13 Timelines Forecast — AI 2027 [REVIEWED] Forecast note: SC definition; method outputs (time-horizon vs benchmarks+gaps); disclosed edits + interpolation bug (~9-month shift)
2026-02-13 AI Futures Model — Forecasts (Daniel 01-26-26) [REVIEWED] Dashboard snapshot: AC ATC percentiles (Daniel vs Eli) + changelog bug-fix deltas and ATC adjustments
2026-02-13 Marcus: “The AI 2027 Scenario: How realistic is it?” [REVIEWED] External critique: narrative/probability hygiene; “fiction vs science”; potential political backfire mechanism
2026-02-13 AI 2027 review research (ChatGPT transcript) [REVIEWED] Secondary memo: replication checklist + circularity flag for uplift resolution
2026-02-13 AI 2027 / AI Futures Model verification memo (Claude transcript) [REVIEWED] Secondary memo: model-revision framing; treat external numeric claims as unverified without primary citations
2026-02-13 AI 2027 Predictions and Primary Resources (independent scorecard) [REVIEWED] Internal pace scorecard (~0.70×): inputs ahead, autonomy/productivity behind; sensitivity-banded estimate
2026-02-13 Bostrom: Optimal Timing for Superintelligence [REVIEWED] Working paper: person-affecting “optimal timing” models often imply modest delays; suggests “swift to harbor, slow to berth” (reach capability quickly, consider brief post-capability pause)
2026-02-11 Reality Check framework: dev velocity + codebase size [REVIEWED] Repo-churn + scc/COCOMO snapshot at v0.3.0; upper-tail feasibility baseline for AI-assisted dev velocity (non-causal)
2026-02-11 Yegge: The AI Vampire [REVIEWED] Medium post framing AI burnout as value-capture/extraction; warns about outlier narratives; proposes a shorter intense workday
2026-02-11 HBR: AI doesn’t reduce work, it intensifies it [REVIEWED] Field study at ~200-person tech company: task expansion, blurred boundaries, multitasking; proposes “AI practice” norms
2026-02-11 Simon Willison: AI intensifies work [REVIEWED] Practitioner reflection: “exhausting productivity,” rapid depletion, and “one more prompt” sleep-disruption anecdotes
2026-02-11 TechCrunch: early signs of burnout among AI embracers [REVIEWED] Journalism synthesis: “burnout machine” framing; highlights rebound/expectation creep; cites HBR/METR/NBER
2026-02-11 METR: AI impact on experienced OSS dev productivity [REVIEWED] RCT-like study: AI allowed → ~19% slower on real OSS issues; large perception vs reality gap
2026-02-11 NBER WP 33777: LLMs, small labor market effects [REVIEWED] Denmark surveys+admin records: modest time savings (~3% typical) + null earnings/hours effects; task reallocation/new tasks
2026-02-11 fast.ai: “dark flow” and vibe coding [REVIEWED] “Dark flow” lens: rapid feedback + delayed-quality signals; miscalibration and maintainability risk hypotheses
2026-02-11 HN discussion: AI intensifies work (46945755) [REVIEWED] Comment-thread snapshot: Jevons-paradox framing, task expansion/review burden anecdotes, and automation-transition debate
2026-02-11 Thread: AI prompting as dopamine loop (@aakashgupta) [REVIEWED] Social framing: reinforcement loop and “one prompt away” stopping-point erosion (speculative quantifiers flagged)
2026-02-11 Thread: “automated coding, not software engineering” (@math_rachel) [REVIEWED] Social hypothesis: agents generate code but not modular abstractions; possible downstream maintenance/burnout burden
2026-02-09 StrongDM “Software Factory” (main essay) [REVIEWED] Vendor manifesto: non-interactive “software factory” posture; scenarios as governance; DTU (behavioral twins) concept; token-fuel heuristic
2026-02-09 StrongDM “Software Factory” (principles) [REVIEWED] Doctrine: seed→validation harness→feedback loop until holdout scenarios pass stably; “frontier engineering” as representation work
2026-02-09 Shapiro: five levels to the “Dark Factory” [REVIEWED] Heuristic maturity model (0–5) for AI coding; role shift coder→reviewer→PM/spec; “technical deflation” framing
2026-02-09 Willison: StrongDM software without looking at code [REVIEWED] Makes “Dark Factory” concrete; highlights spec-only strongdm/attractor artifact + holdout-scenario framing
2026-02-09 Anthropic: harnesses for long-running agents [REVIEWED] Harness pattern: initializer + iterative coding sessions + durable artifacts (progress logs, git checkpoints) + end-to-end testing discipline
2026-02-09 Schillace: compounding teams [REVIEWED] “Compounding” org pattern: bespoke frameworks + filesystem/git artifacts + recursive tool building; warns against mixing manual+AI styles in one repo
2026-02-09 DoltHub: a day in Gas Town [REVIEWED] Field report on Gas Town “limitless mode”; highlights merge/guardrail failure modes and cost; suggests Dolt as memory substrate
2026-02-09 Virtuslab: Beads review (“give AI memory”) [REVIEWED] Review of Beads as agent memory/task system; agent-UX heuristics (structured output, deterministic thick logic, idempotency)
2026-02-09 Yegge: “Revenge of the junior developer” [REVIEWED] Early “waves” model: agents→clusters→fleets; token-opex framing; “junior revenge” labor hypothesis under rising spend
2026-02-09 Thread: “Windows is awful” + caste-discrimination claims at Microsoft (Thread Reader) [REVIEWED] Low-rigor, high-bias thread: separates general caste-discrimination reporting/cases (Cisco verified) from uncorroborated Microsoft+“Python vs C++” allegations; X embeds inaccessible
2026-02-07 Superlinear Multi-Step Attention (arXiv 2601.18401v1) [REVIEWED] Preprint: structural non-exclusion framing; N=2 O(L^(3/2)) scaling; self-reported B200 throughput + NIAH results
2026-02-07 Superlinear (GitHub repo @ df9b2ef) [REVIEWED] Repo audit: Triton kernels + OpenAI-style server; explicit sessions/snapshots for KV cache reuse; memory budgeting cross-checked
2026-02-07 Superlinear-Exp-v0.1 (HF model release) [REVIEWED] Weights+remote code audit: trust_remote_code; ~59GiB weights; config implies ~6GB/1M tokens KV; license + patent notice
2026-02-07 Semantica (GitHub repo @ b4cfb6d) [REVIEWED] Repo audit: claims vs code/tests; maturity, strengths/weaknesses, and gaps
2026-02-06 Time Horizon 1.1 (METR) [REVIEWED] METR methodology update: suite 170→228, long tasks 14→31, Vivaria→Inspect; updated slope estimates + protocol/task-distribution sensitivity
2026-02-06 How Does Time Horizon Vary Across Domains? (METR) [REVIEWED] Cross-domain time-horizon comparisons: coding/math/QA fast; GUI agents (OSWorld) much shorter; proxy-based caveats + soundness conditions
2026-02-06 Measuring AI Ability to Complete Long Tasks (METR blog) [REVIEWED] Introduces time-horizon metric; reports ~7-month doubling with wide CIs; includes partial reproduction checks via released run data
2026-02-06 Introducing GPT-5.3-Codex [REVIEWED] OpenAI release post: agentic coding + “computer work” framing; benchmark table; Terminal-Bench leaderboard cross-check; cyber capability classification + trusted-access program
2026-02-06 GPT-5.3-Codex System Card [REVIEWED] System card: Codex sandboxing + network controls; destructive-action avoidance mitigation; Preparedness “High cyber capability” framing with layered safeguards + TAC; Apollo sandbagging/sabotage summary; residual risk notes
2026-02-06 Claude Opus 4.6 [REVIEWED] Anthropic release post: 1M context (beta), compaction/effort controls, pricing tiers; benchmark claims vs Terminal-Bench leaderboard; GDPval-AA Elo cross-check
2026-02-06 System Card: Claude Opus 4.6 [REVIEWED] System card: training data cutoff (May 2025); adaptive thinking + effort levels; benchmark methodology; RSP/ASL-3 deployment determination; alignment key findings incl overly agentic GUI computer-use and “suspicious side tasks”; external testing (AISI/Apollo/Andon)
2026-02-06 AI 2027 [REVIEWED] Scenario + forecast pages: “superhuman coder” by 2027 plausible; ~1-year superhuman-coder→ASI median; compute stock 10× to 100M H100e; theft/race dynamics
2026-02-01 Memorandum of Strategy: Project Mirror (Arclight) [DRAFT] Proposal: anti-sycophancy AI “mentor” optimized for graduation; psychometric middleware + lockout rules; limited empirical grounding
2026-02-01 Arclight Operating System (AOS): Internal Codex [DRAFT] Constraint-based institution “OS”: agency + endowment + founder-decay; argues kernel must stay hidden; heavy on non-operational metaphors
2026-02-01 Reality Check (framework README) [REVIEWED] Repo self-description; verified CLI+DB primitives; notes rigor-v1 validator + 401 tests
2026-02-01 Reality Check Knowledge Base (dataset README) [REVIEWED] Dataset snapshot and structure; verified DB counts and layout; flags YAML exports vs DB truth
2026-02-01 Minnesota standoff with Trump administration stokes fears of civil war (The Hill) [REVIEWED] Civil-war escalation framing; links Good/Pretti killings to federalism/legal conflict (SCOTUS Guard limits) and CERL simulation; flags under-verified deployment-scale claims
2026-01-31 How the World Sees America, With Adam Tooze (NYT) [REVIEWED] Davos 2026 “rupture” framing; Greenland tariffs as rupture datapoint; China scale + climate constraints; skepticism of “U.S.→China hegemony transition” story
2026-01-31 DHS retreats from the claim that the agents who killed Alex Pretti faced a 'violent riot' (Reason) [DRAFT] “Violent riot” framing vs OPR-described yelling/whistles; Bovino “rights don’t count” quote; credibility argument via prior Chicago misrepresentation finding
2026-01-30 These Trump voters 'formed a suicide pact' and Republicans are panicking: ex-GOP operative (Raw Story) [DRAFT] Opinion recap of Rick Wilson on farm distress from tariffs + immigration enforcement; verified “444 farming counties” typology + high foreign-born crop labor share; key numeric forecasts uncorroborated
2026-01-27 Special Report: China sets new records in air-sea operations around Taiwan (Janes) [DRAFT] Verification source: Taiwan MND-attributed ADIZ annual totals (2023 vs 2024) used as pressure metric
2026-01-27 Military Implications of PLA Aircraft Incursions in Taiwan’s Airspace 2024 (Jamestown) [DRAFT] Verification source: definition-sensitive median-line day-count metric (62 vs 209 “over half” days)
2026-01-27 Energy supplies sufficient, ministry says (Taipei Times) [DRAFT] Verification source: MOEA-stated natural-gas reserves ~10–11 days of consumption (2022-08)
2026-01-27 Stable Supply of Natural Gas (Taiwan MOEA Energy Administration) [DRAFT] Verification source: policy target schedule for natural-gas “security stockpile required” (14 days by 2027)
2026-01-27 China's January 2026 military purge reshapes Taiwan calculus (Claude deepresearch memo) [DRAFT] AI-generated synthesis memo: blockade/quarantine framing, Taiwan energy vulnerability, and norms/credibility context
2026-01-26 2026 United States intervention in Venezuela (Wikipedia) [DRAFT] Context source: Venezuela intervention capture event + UN/legal criticism as precedent/analogy
2026-01-26 How a purge of China’s military leadership could impact the army and the future of Taiwan (AP) [DRAFT] Explainer: announced investigations, purge scale, and competing views on short- vs long-term Taiwan risk
2026-01-26 Taiwan monitoring “abnormal” China military leadership changes (Reuters) [DRAFT] Taiwan official posture: monitor abnormal PRC changes; don’t lower guard; notes near-daily PRC ops
2026-01-26 What Xi Jinping’s purge of China’s most senior general reveals (Economist) [DRAFT] Motive space: corruption vs performance vs faction; warns of readiness impacts; cites DoD assessments
2026-01-26 Xi’s Purge of Top General Spurs Questions on Taiwan, Succession (Bloomberg) [DRAFT] Speed/urgency and succession speculation; skeptical of nuclear-leak claim; Taiwan timing implications
2026-01-26 Xi takes sole operational control of army as China probes military leaders (FT) [DRAFT] CCP “discipline” language caution; “sole operational control” framing; ties to 2027 capability target discourse
2026-01-26 Xi’s Purge of China’s Military Brings Its Top General Down (NYT) [DRAFT] Xi distrust hypothesis; public investigation implies downfall; suggests multi-year rebuild lowers short-term invasion odds
2026-01-26 Xi's purge paints a picture of power and control (Sky News) [DRAFT] Signal-reading: speed + PLA Daily political framing; captures coup/treason rumor ecology
2026-01-26 Military purge gives Xi total control of Chinese army (Telegraph) [DRAFT] Strong-form analyst forecasts (“Taiwan delayed for years”) and purge-tempo-as-strength hypothesis
2026-01-26 China’s Top General Accused of Giving Nuclear Secrets to U.S. (WSJ) [DRAFT] Unnamed-source briefing allegations: nuclear leak + bribes + Shenyang task force; readiness-vacuum risk
2026-01-26 The Chinese Spy Machine Infiltrating Taiwan’s Military (WSJ) [DRAFT] Taiwan insider-threat axis: espionage cases and recruitment methods; relevance to contingency risk
2026-01-26 PLA purges intensify: Zhang Youxia and Liu Zhenli under investigation (Sinocism) [DRAFT] PLA Daily excerpt emphasizes authority/cliques; SCMP-sourced internal briefings/detention claims (provisional)
2026-01-26 Hot takes/threads on Zhang Youxia purge and Taiwan timing (X compilation) [DRAFT] Rumor-space map: invasion-delay, info-warfare, coup claims; extracted as low-evidence hypotheses
2026-01-26 Physician Declaration on Pretti Shooting (Doc. 109) [DRAFT] Primary court filing: physician declaration describing post-shooting medical response + reported wound locations (agency attribution in declaration is disputed)
2026-01-26 Alex Pretti Family Statement (Common Dreams) [REVIEWED] Family response to CBP shooting; video evidence disputes official narrative (phone vs gun, disarmament before shooting); immediate "domestic terrorist" labeling without evidence
2026-01-25 "Positions ICE has taken" (BlueBadger2600) [REVIEWED] Legal/policy fact-check of ICE authority claims (home entry, detention, counsel, courts, third-country removals) + Minnesota incidents (Good/Pretti, 5yo detention, National Guard, Vance immunity)
2026-01-25 Inference Cost Analysis & Projections (lhl) [DRAFT] Frontier inference unit economics: caching + routing dominate; reasoning/long-context are hidden cost multipliers; includes cache-heavy case study
2026-01-24 Gas Town’s Agent Patterns (Appleton) [REVIEWED] Design-fiction read of Gas Town; bottleneck shifts to design/verification; orchestration patterns + “code distance” axis
2026-01-24 JP-TL-Bench paper (Lin & Lensenmayer 2026) [REVIEWED] Anchored pairwise LLM-judge benchmark for JA↔EN translation; Bradley–Terry aggregation + LT scoring; protocol dependence noted
2026-01-24 JP-TL-Bench (blog) [REVIEWED] Release post motivating anchored pairwise evaluation vs metric compression for high-end JA↔EN translation
2026-01-24 JP-TL-Bench Directional Analysis [REVIEWED] Direction and difficulty slices show large asymmetries (e.g., Llama 3.1 8B); qualitative examples
2026-01-24 JP-TL-Bench repo [REVIEWED] Open artifacts: Base Set v1.0 snapshot, prompts, scoring code; reproducibility and bias considerations
2026-01-24 Shisa V2 [REVIEWED] Model family release context; claims about Japanese eval gaps and new translation benchmark motivation
2026-01-24 Shisa V2 405B [REVIEWED] Model release positioning; internal eval suite framing; translation evaluation rationale
2026-01-24 Shisa V2.1 [REVIEWED] Dataset/recipe refresh claims; context for Japanese-specific improvements (incl. translation/politeness)
2026-01-24 Liquid AI Hackathon Tokyo repo [REVIEWED] COMET-centered llm-jp-eval MT workflow; prompt asymmetry and metric clustering cautions
2026-01-24 LiquidAI LFM2-350M-ENJP-MT model card [REVIEWED] Vendor description of a small bidirectional EN↔JA MT checkpoint for low-latency use
2026-01-24 TranslateGemma Technical Report [REVIEWED] SFT+RL translation specialization of Gemma 3; WMT metrics + MQM human eval; en→ja gains reported
2026-01-24 TranslateGemma blog [REVIEWED] Announcement summary: model sizes, language coverage, and metric-based gains
2026-01-24 FinanceBench (Islam et al. 2023) [REVIEWED] Benchmark baseline: 10,231 Qs; human-eval sample shows high failure rates for GPT-4-Turbo+retrieval; oracle ~85%
2026-01-24 @_avichawla thread on PageIndex [REVIEWED] Social summary of “vectorless RAG”; highlights similarity≠relevance + structure-aware retrieval; performance claim flagged as protocol-dependent
2026-01-24 VectifyAI PageIndex repo [REVIEWED] Repo review: hierarchical tree index + LLM traversal (no vector DB); strong adoption but limited CI/tests; “98.7% FinanceBench” is vendor-reported
2026-01-24 VectifyAI Mafin 2.5 FinanceBench results repo [REVIEWED] Eval artifacts: claims 98.7% via permissive LLM-judge; not directly comparable to FinanceBench human-eval baselines
2026-01-24 StrataLens AI repo [REVIEWED] Single-maintainer hybrid RAG (pgvector+rerank+section routing); claims 85% on FinanceBench subset; repo license file missing
2026-01-23 Anthropic "Claude's Constitution" [REVIEWED] Values/behavior spec: “final constitutional authority”; hard constraints (WMD uplift, critical infrastructure attacks, malware, CSAM); virtue-ethics + model-welfare framing
2026-01-22 Stross "The pivot" (Part 1) [REVIEWED] 2025 inflection-point thesis: energy transition + climate/food risk + post-Moore’s Law tech/finance
2026-01-22 OpenAI "Value Intelligence" [REVIEWED] Compute + ARR scaling claims; compute-scarcity flywheel; tiered monetization framing
2026-01-21 Rogan & Musk "JRE #2404 — Elon Musk" [REVIEWED] Transcript excerpt: truth-seeking alignment, Gemini ImageGen controversy, and objective mis-specification risks
2026-01-21 Chatterjee et al. "ANZ Copilot Study" [REVIEWED] ANZ A/B study: Copilot group completes coding challenges ~42.36% faster
2026-01-21 Peng et al. "GitHub Copilot Productivity" [REVIEWED] Controlled trial: Copilot users complete a coding task ~55.8% faster
2026-01-21 Carney "Middle Powers & Rupture" [REVIEWED] WEF Davos speech on coalition strategy vs. great power coercion
2026-01-18 Doctorow "Reverse Centaur" [REVIEWED] AI bubble critique + "reverse centaur" deployment patterns
2026-01-18 Perera "China's Trillion-Dollar Illusion" [REVIEWED] "Demand-destruction surplus" mechanism analysis
2026-01-18 Ronacher "Agent Psychosis" [REVIEWED] Maintainer-centric analysis of agentic-coding addiction/slop loops
2026-01-16 Teortaxes "Greenland Endgame" [REVIEWED] GOV/RESOURCE/TRANS claims extraction with fact-checks

Claim Domains

Domain Description Claims
TECH Technology & AI 245
LABOR Labor & Employment 47
ECON Economics & Markets 110
GOV Governance & Policy 303
SOC Social Dynamics 54
TRANS Transition Dynamics 60
RESOURCE Resource Constraints 20
GEO Geopolitics 63
INST Institutions & Organizations 171
RISK Risk Assessment 89
META Framework & Methodology 115

Repository Structure

realitycheck-data/
├── README.md                 # This file
├── .realitycheck.yaml        # Framework configuration
│
├── data/
│   └── realitycheck.lance/   # LanceDB database (git-lfs)
│
├── claims/
│   ├── README.md             # Claim statistics (auto-generated)
│   ├── registry.yaml         # Exported claim registry
│   └── chains/               # Argument chain analyses
│
├── reference/
│   ├── sources.yaml          # Source bibliography
│   ├── primary/              # Primary source materials
│   ├── transcripts/          # AI-assisted analysis transcripts
│   ├── articles/             # Article references
│   └── data/                 # Data source references
│
├── analysis/
│   ├── sources/              # Individual source analyses
│   ├── syntheses/            # Cross-source syntheses
│   └── meta/                 # Framework meta-analyses
│
├── tracking/
│   ├── predictions.md        # Prediction tracking
│   ├── dashboards/           # Indicator dashboards
│   └── updates/              # Prediction updates
│
├── scenarios/                # Scenario matrices
│
└── inbox/                    # Work tracking
    ├── to-catalog/
    ├── to-analyze/
    └── in-progress/

Using This Knowledge Base

Prerequisites

Install the Reality Check framework:

# Clone the framework
git clone https://github.com/lhl/realitycheck.git
cd realitycheck
uv sync  # Install dependencies

Validate Data Integrity

export REALITYCHECK_DATA=/path/to/realitycheck-data/data/realitycheck.lance
uv run python scripts/validate.py --mode db

Query Claims

# Search claims
uv run python scripts/db.py search "automation labor"

# List claims by domain
uv run python scripts/db.py claim list --domain LABOR

# Get specific claim
uv run python scripts/db.py claim get TECH-2026-001

Export Data

# Export claims to YAML
uv run python scripts/export.py yaml claims -o claims/registry.yaml

# Export summary to Markdown
uv run python scripts/export.py md summary -o claims/README.md

Research Questions

  1. Transition Dynamics: How do economies transition to post-scarcity?
  2. Value Theory: What happens to value when marginal cost → 0?
  3. Distribution: Who owns AI/automation means of production?
  4. Labor: What happens to human work and meaning?
  5. Governance: What structures handle coordination in post-labor world?
  6. Timeline: How plausible are 2035-2045 predictions?
  7. Bottlenecks: Who controls compute, energy, chips?

License

  • Framework: Reality Check - Apache 2.0 License
  • Analyses: Cite appropriately

Last updated: 2026-09-16

About

Reality Check knowledge base - unified analysis of claims across technology, economics, labor, and governance domains

Resources

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages