Skip to content

feat(ai): admin-only paid OpenRouter + refresh model catalog & pricing - #1473

Merged
2witstudios merged 6 commits into
masterfrom
pu/openrouter
Jun 2, 2026
Merged

2witstudios merged 6 commits into
masterfrom
pu/openrouter

Conversation

@2witstudios

@2witstudios 2witstudios commented Jun 1, 2026 •

Copy link
Copy Markdown
Owner

Summary

Exposes the paid OpenRouter provider to global admins only (gated on role === 'admin', independent of subscription tier), refreshes the model catalog against the live OpenRouter API, and keeps the metering price table in sync.

Admin-only gating (role-based, not tier)

  • New ADMIN_ONLY_PROVIDERS + isAdminOnlyProvider() in ai-providers-config.ts as the single source of truth.
  • GET/PATCH /api/ai/settings: provider visibility is now role-aware — admins see & can select paid OpenRouter; everyone else (including pro/business) gets it masked → 503 on PATCH. Admins bypass the Pro-subscription gate.
  • requiresProSubscription() gains an isAdmin short-circuit.
  • /api/ai/chat enforces admin-only at generation time (defense-in-depth: handles a demoted ex-admin whose stored selection is still paid OpenRouter).
  • No client changes — the model selector already renders whatever /api/ai/settings reports as available.

A paid-tier non-admin gets no access; an admin on the free tier does. Admins still count against the standard/pro daily rate-limit quota.

Catalog refresh (hit the live API — 343 models)

Diffed the static list against the live OpenRouter catalog:

  • Added current tool-capable chat flagships: Claude Opus 4.8 (+ Fast), GPT-5.2 Pro/Chat, GPT-5.1 Chat, GPT-5.1 Codex Max, GPT-5 Pro/Codex, o3, o3-pro, o4-mini, GPT-4.1/-mini, GLM 5.1, GLM 5 Turbo, GLM 4.6, GLM 4.7 Flash, Kimi K2.6, Kimi K2 Thinking, MiniMax M2.1, MiniMax M3, Qwen3.7 Max, Mistral Large 3, Mistral Medium 3, Devstral 2, Grok Build 0.1.
  • Removed 12 stale IDs that 404 on OpenRouter (claude-3.5-sonnet, gemini-2.0-pro, grok-4, grok-4-fast, llama-3.1-405b-instruct, devstral-medium/small, gpt-5.2-mini/nano, jamba-mini-1.7, inception/mercury, gemini-2.5-flash-lite-preview-06-17).
  • Result: 120 models, all confirmed live (zero stale). Image/audio/:free/legacy variants were intentionally excluded (the selector is for tool-using chat).

Metering / pricing (the important bit 💸)

calculateCost() falls back to default: { input: 0, output: 0 }, so any model missing from AI_PRICING is metered as free. Added per-1M prices (pulled from the live OpenRouter API) for every newly added model, achieving full catalog↔pricing parity — and fixed a pre-existing gap where writer/palmyra-x5 was unpriced.

Verification

  • typecheck: clean across all changed source files (after building @pagespace/db + @pagespace/lib).
  • Tests: 63 pass across settings route / rate-limit-middleware / ai-providers-config; 62 ai-monitoring tests pass.
  • Catalog↔pricing parity verified programmatically (0 unpriced paid models).

Provider surfaces covered

Both per-user model-selection paths go through the same gated ProviderModelSelector (chat footer + global assistant via ChatInput → InputFooter), so admins see paid OpenRouter in both and non-admins never do. Backend enforcement (admin-only guard + admin subscription bypass) is applied on both the page-AI chat route (/api/ai/chat) and the per-user global assistant route (/api/ai/global/[id]/messages).

Genuinely out of scope: shared page-agent (AI Chat page) configuration used by other drive members — that is a shared surface with different semantics and is not part of this per-user admin-gating change.

Gate the paid OpenRouter provider to global admins (role === 'admin'),
independent of subscription tier — a paid-tier non-admin gets no access,
an admin on the free tier does.

- Add ADMIN_ONLY_PROVIDERS + isAdminOnlyProvider() as the single source of truth
- /api/ai/settings GET/PATCH: role-aware provider visibility; admins bypass the
  Pro-subscription gate (still counted against daily quotas)
- requiresProSubscription() gains an isAdmin short-circuit
- /api/ai/chat enforces admin-only at generation time (defense-in-depth for a
  demoted ex-admin's stored selection)

Refresh the paid OpenRouter catalog against the live API (343 models): add
current tool-capable flagships (Claude Opus 4.8, GPT-5.2 Pro/Chat, GLM 5.1,
Kimi K2.6, MiniMax M3, Qwen3.7 Max, etc.) and drop 12 stale IDs that 404 on
OpenRouter. Now 120 models, all confirmed live.

Keep the AI_PRICING metering table in sync: add per-1M prices (from the live
OpenRouter API) for every newly added model — the cost fallback is $0, so a
missing entry would meter the model as free. Achieves full catalog↔pricing
parity (also fills a pre-existing gap for writer/palmyra-x5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
pagespace-marketing Ready Ready Preview, Comment Jun 2, 2026 12:06am

@coderabbitai

coderabbitai Bot commented Jun 1, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@2witstudios, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 22 minutes and 34 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: fb2eda88-5874-45fe-a730-3e57949fcada

📥 Commits

Reviewing files that changed from the base of the PR and between 0d300c8 and 8b0efa7.

📒 Files selected for processing (7)
  • apps/web/src/app/api/ai/chat/__tests__/stream-socket-events.test.ts
  • apps/web/src/app/api/ai/chat/route.ts
  • apps/web/src/app/api/ai/global/[id]/messages/__tests__/stream-socket-events.test.ts
  • apps/web/src/app/api/ai/global/[id]/messages/route.ts
  • apps/web/src/lib/subscription/rate-limit-middleware.ts
  • packages/lib/src/monitoring/__tests__/ai-monitoring.test.ts
  • packages/lib/src/monitoring/ai-monitoring.ts
📝 Walkthrough

Walkthrough

This PR establishes admin-only gating for the paid OpenRouter AI provider and refreshes the model/pricing catalog. It exports ADMIN_ONLY_PROVIDERS and an isAdminOnlyProvider() helper, modifies requiresProSubscription() to bypass Pro checks for admins, enforces admin-only restrictions in chat and settings routes via role-aware provider visibility and early rejections, and updates OpenRouter and other provider model mappings with new Claude, GPT, and specialized model entries.

Changes

Admin-only provider gating and enforcement

Layer / File(s) Summary
Admin-only provider definitions
apps/web/src/lib/ai/core/ai-providers-config.ts, apps/web/src/lib/ai/core/__tests__/ai-providers-config.test.ts
Exports ADMIN_ONLY_PROVIDERS set (containing openrouter) and isAdminOnlyProvider() predicate; tests verify openrouter is marked admin-only while free/standard providers are not.
Admin bypass for subscription gating
apps/web/src/lib/subscription/rate-limit-middleware.ts, apps/web/src/lib/subscription/__tests__/rate-limit-middleware.test.ts
requiresProSubscription() gains isAdmin parameter (default false) with early return bypass when true; test verifies admins on free tier can use paid models without subscription.
Chat route admin-only provider checks
apps/web/src/app/api/ai/chat/route.ts, apps/web/src/app/api/ai/chat/__tests__/stream-socket-events.test.ts
Chat POST and PATCH handlers compute isAdminUser from auth role, reject non-admins selecting admin-only providers with 503, and pass admin flag to requiresProSubscription(); test mock updated to include isAdminOnlyProvider.
Settings route role-based provider visibility
apps/web/src/app/api/ai/settings/route.ts, apps/web/src/app/api/ai/settings/__tests__/route.test.ts
New visibleProvidersFor(role) helper exposes ADMIN_ONLY_PROVIDERS to admins only; GET and PATCH use role-aware visibility for provider filtering and selection validation; subscription checks pass auth.role === 'admin' bypass flag; tests cover admin/non-admin visibility and subscription bypass scenarios.

AI model catalog refresh

Layer / File(s) Summary
Provider model catalog updates
apps/web/src/lib/ai/core/ai-providers-config.ts
Adds Claude Opus 4.8 (standard/fast), expands OpenAI GPT-5.x/o-series, and updates Gemma, Llama, Mistral, GLM/z-ai, Qwen, Moonshot, MiniMax, xAI, and AI21 models in OpenRouter paid provider mappings.
Pricing catalog updates
packages/lib/src/monitoring/ai-monitoring.ts
Adds per-1M token pricing for new Anthropic Claude Opus 4.8, OpenAI GPT-5.2/5.1/o3/o4-mini/GPT-4.1, Mistral large/medium variants, z-ai GLM, Qwen, Moonshot Kimi, MiniMax, and xAI Grok entries; comments note delisted Grok models retained for metering.

Sequence Diagram

sequenceDiagram
  participant Client
  participant SettingsRoute
  participant visibleProvidersFor
  participant requiresProSubscription
  Client->>SettingsRoute: GET /api/ai/settings<br/>(with auth.role)
  SettingsRoute->>visibleProvidersFor: role
  visibleProvidersFor-->>SettingsRoute: user providers + admin providers if admin
  SettingsRoute-->>Client: available providers filtered by visibility
  Client->>SettingsRoute: PATCH /api/ai/settings<br/>(select provider, with auth.role)
  SettingsRoute->>visibleProvidersFor: role
  visibleProvidersFor-->>SettingsRoute: visible provider set
  alt provider in visible set
    SettingsRoute->>requiresProSubscription: provider, model, tier, isAdmin
    alt isAdmin or tier is Pro
      requiresProSubscription-->>SettingsRoute: false (bypass)
      SettingsRoute-->>Client: 200 OK, setting updated
    else free tier and not admin
      requiresProSubscription-->>SettingsRoute: true (required)
      SettingsRoute-->>Client: 503 Subscription required
    end
  else provider not in visible set
    SettingsRoute-->>Client: 403 Provider not available
  end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • 2witstudios/PageSpace#629: Both PRs modify apps/web/src/lib/subscription/rate-limit-middleware.ts—the main PR extends requiresProSubscription to add an isAdmin bypass flag and wire it from the AI routes, while the retrieved PR refactors requiresProSubscription to gate based on getPageSpaceModelTier.
  • 2witstudios/PageSpace#1397: Both PRs modify the AI settings API route's provider-visibility/selection logic to change which providers are allowed for different user contexts (admin-only gating vs. enabling openrouter_free).
  • 2witstudios/PageSpace#795: Both PRs modify the Pro/subscription gate in rate-limit-middleware.ts—the main PR adds an isAdmin bypass flag while the retrieved PR switches the early-return condition from on-prem to billing-enabled behavior.

Poem

🐰 Hops through the gates with admin might,
OpenRouter shines for those with right,
New models bloom across the lands,
Pro subscriptions—admins skip these strands,
The catalog grows, a feast so bright! 🎯✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately describes the main changes: introducing admin-only gating for paid OpenRouter and refreshing the model catalog with pricing updates.
Docstring Coverage ✅ Passed Docstring coverage is 83.33% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pu/openrouter

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b256bfbc6c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

'x-ai/grok-4.20-multi-agent': { input: 2.00, output: 6.00 },
'x-ai/grok-4-fast': { input: 0.20, output: 0.50 },
'x-ai/grok-4': { input: 3.00, output: 15.00 },
'x-ai/grok-build-0.1': { input: 1.00, output: 2.00 },

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve pricing for existing Grok 4 selections

When an admin has an existing saved OpenRouter selection or passes selectedModel for x-ai/grok-4/x-ai/grok-4-fast, createAIProvider still sends that model ID to OpenRouter without consulting AI_PROVIDERS, but this replacement removes their AI_PRICING entries. trackAIUsage then calls calculateCost, which falls back to AI_PRICING.default (0), so those paid calls are logged as free and billing/monitoring underreports costs.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — fixed in 9776f5e (now merged into the branch). Pricing is now treated as a superset of the selectable catalog: x-ai/grok-4 and x-ai/grok-4-fast were removed from the picker but their AI_PRICING entries are retained (with a comment explaining why), so any lingering saved selection still meters at the correct rate instead of falling back to the $0 default. This also matches how the other delisted models were handled — they were only removed from the catalog, never from pricing.

2witstudios and others added 3 commits June 1, 2026 18:22
Addresses Codex review (P2): removing x-ai/grok-4 and x-ai/grok-4-fast from
AI_PRICING would meter any lingering saved selection at the $0 default fallback.
Pricing is now a superset of the selectable catalog — entries for delisted models
are retained for correct cost tracking while the models stay out of the picker.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts:
#	apps/web/src/lib/subscription/__tests__/rate-limit-middleware.test.ts
The chat route now calls isAdminOnlyProvider() in the hot path. The
stream-socket-events test mocks ai-providers-config, so the missing export
made the call throw → route 500'd before createStreamLifecycle, failing all
13 lifecycle assertions. Add isAdminOnlyProvider to the mock (returns false).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/web/src/app/api/ai/chat/route.ts`:
- Around line 433-440: The code currently uses
createSubscriptionRequiredResponse() when isAdminOnlyProvider(currentProvider)
blocks non-admins, which returns a misleading 403 message about subscription;
change this to return a dedicated admin-restriction response (e.g.,
createAdminRestrictedResponse()) or directly return the same error shape/message
used by validateProviderModel ("This provider is restricted to administrators")
so the client sees the correct reason; update the call site in route.ts (the
block using isAdminOnlyProvider) to call the new helper instead of
createSubscriptionRequiredResponse(), and add the helper function (or adjust the
existing response helpers) to centralize the 403 + correct error message for
admin-only provider rejections.

In `@packages/lib/src/monitoring/ai-monitoring.ts`:
- Around line 20-21: The new AI_PRICING entries lack corresponding
MODEL_CONTEXT_WINDOWS entries so getContextWindow() falls back to the default;
update MODEL_CONTEXT_WINDOWS to add explicit context-window sizes keyed by each
new priced model id (the exact keys from AI_PRICING such as
'anthropic/claude-opus-4.8', 'anthropic/claude-opus-4.8-fast', and the other IDs
called out in the comment) so they map to the correct token/window values, or
alternatively add a parity guard test that asserts every key in AI_PRICING has a
matching key in MODEL_CONTEXT_WINDOWS (use the existing AI_PRICING and
MODEL_CONTEXT_WINDOWS symbols and the getContextWindow() usage to locate where
to add entries/tests).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 8f719525-f2e5-4d6d-a05b-639c0a224a3d

📥 Commits

Reviewing files that changed from the base of the PR and between edf1233 and 0d300c8.

📒 Files selected for processing (9)
  • apps/web/src/app/api/ai/chat/__tests__/stream-socket-events.test.ts
  • apps/web/src/app/api/ai/chat/route.ts
  • apps/web/src/app/api/ai/settings/__tests__/route.test.ts
  • apps/web/src/app/api/ai/settings/route.ts
  • apps/web/src/lib/ai/core/__tests__/ai-providers-config.test.ts
  • apps/web/src/lib/ai/core/ai-providers-config.ts
  • apps/web/src/lib/subscription/__tests__/rate-limit-middleware.test.ts
  • apps/web/src/lib/subscription/rate-limit-middleware.ts
  • packages/lib/src/monitoring/ai-monitoring.ts

Comment thread apps/web/src/app/api/ai/chat/route.ts
Comment thread packages/lib/src/monitoring/ai-monitoring.ts
Addresses CodeRabbit review:
- Minor: the chat route's admin-only block returned createSubscriptionRequiredResponse()
  ("requires a Pro/Business subscription"), which is misleading — the block is on
  role, not tier. Add createAdminRestrictedResponse() (403, "restricted to
  administrators") and use it. Mock it in the stream-socket-events test.
- Major: every newly priced model lacked a MODEL_CONTEXT_WINDOWS entry, so
  getContextWindow() fell back to the 200k default (risking truncation / provider
  context-limit errors). Add accurate context windows (from the live OpenRouter API)
  for all 26 new models, achieving full AI_PRICING<->MODEL_CONTEXT_WINDOWS parity.
- Add a parity guard test so future pricing additions can't drift from context windows.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…route

Global assistants are per-user (not shared), so they're a primary surface for the
admin-only paid OpenRouter option — and they already expose it via the footer's
gated ProviderModelSelector. Add the matching backend guard to
/api/ai/global/[id]/messages so admin-only providers are blocked for non-admins
(defense-in-depth against a stored selection surviving a role downgrade), mirroring
the page-AI chat route. Admins are unaffected (the route has no subscription gate).

Update the two stream-socket-events test mocks to declare the new exports the route
now consumes (isAdminOnlyProvider, createAdminRestrictedResponse).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@2witstudios
2witstudios merged commit c97147a into master Jun 2, 2026
12 checks passed
2witstudios added a commit that referenced this pull request Jun 2, 2026
@2witstudios
2witstudios deleted the pu/openrouter branch June 2, 2026 03:41
2witstudios added a commit that referenced this pull request Jun 3, 2026
…econcile, accounting + e2e (#1484)

* fix(billing): meter all AI provider calls and bill resolved model names

Close revenue leaks where AI provider calls bypassed credit metering or
recorded $0 cost. Billing only happens via AIMonitoring.trackUsage →
consumeCredits with a real userId and a model present in AI_PRICING.

Unmetered call sites — add trackUsage (real userId was already in scope):
- ask_agent (agent-communication-tools): the largest leak — a user-triggerable
  tool loop of up to stepCountIs(20) round-trips, never billed; also reached by
  every channel @mention of an agent. Returns full ProviderResult from
  getConfiguredModel and meters response.totalUsage so every round-trip counts.
- Memory discovery/integration/compaction: run per active user on memory cron,
  on the expensive pro/glm-5 tier; discovery fires 3 passes/run.
- Zoom extract-action-items and generate-summary: per webhook.

Mis-metered ($0) call sites — track the resolved providerResult.modelName
instead of the raw stored model (PageSpace tier aliases 'standard'/'pro' and
the unpriced default 'glm-4.5-air' all hashed to AI_PRICING.default = $0):
- /api/v1/chat/completions (was page.aiModel ?? 'unknown')
- page-agents/consult (was agent.aiModel || 'glm-4.5-air'); also switch to
  result.totalUsage since it is a stepCountIs(100) tool loop
- /api/ai/chat (was raw currentModel)

Catalog↔pricing drift:
- Add 'glm-4.5-air' to AI_PRICING (0.35/1.55, matching z-ai/glm-4.5-air); it was
  selectable via the glm provider but unpriced, so it metered at $0.

Correctly-metered paths (global assistant, pulse generate/cron, workflow
executor) already used providerResult.modelName and are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(billing): lock in metering for ask_agent and glm-4.5-air pricing

Regression guards for the leak fixes in this PR:
- ai-monitoring: assert PageSpace-tier backend models (glm-4.5-air, glm-4.7,
  glm-5) all price above $0, and that glm-4.5-air bills at its published rate.
  Catches future catalog↔pricing drift that would meter at $0.
- agent-communication-tools: assert ask_agent bills the requesting user against
  the resolved model name (glm-5) using totalUsage (all tool-loop round-trips),
  proving the previously-unmetered path is now metered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): make AI-usage persistence durable and stop $0 ledger churn

Two metering-correctness gaps left open by #1475 (audit-leaks), found in the
prepaid AI-credits audit. Both are pure correctness/safety; no billing-policy
change.

1. Durability (was fire-and-forget). `trackAIUsage` detached the
   writeAiUsage -> consumeCredits chain with `.then()` and returned before it
   settled. Callers `await` trackAIUsage from a stream onFinish / post-response
   handler, but a serverless freeze could drop the detached promise — losing
   BOTH the usage log AND the charge. With no aiUsageLogs row, the reconcile
   cron's orphan sweep has nothing to recover from, so the charge is gone for
   good. Now the chain is awaited, so the write is durable before the request
   returns. Still never throws into the AI request.

2. $0 ledger churn. A free/local model — or a tool-only analytics log with no
   tokens (trackAIToolUsage) — produced amountCents 0, yet consumeCredits still
   opened a balance transaction, took the row lock, and ran a $0 decrement,
   writing a misleading "applied/monthly" ledger row per call. In an agent tool
   loop that serialized N no-op locks on the user's balance row. Now a
   zero-charge call settles the claimed row as 'skipped' without the balance
   transaction. The claim row still exists, so the orphan sweep stays idempotent
   and never re-processes it.

Tests: +1 durability test (asserts consume runs before the awaited trackAIUsage
resolves, no setTimeout flush) and +1 zero-charge test (no transaction, row
marked 'skipped'). credit-consume + ai-monitoring suites: 76 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): extend usage-persistence durability to the tool-call path

Addresses Codex review (P2) on #1476: trackAIToolUsage called trackAIUsage
without returning/awaiting it, so a caller that `await`s trackToolUsage resolved
immediately — the durability guarantee didn't reach tool-analytics logs, and the
writeAiUsage / zero-charge ledger settlement could still be dropped on a
serverless freeze after onFinish.

trackAIToolUsage now RETURNS the trackAIUsage promise (no longer an async wrapper
that discards it), so awaiting it waits for the log to persist. +1 test asserting
the returned promise stays pending until writeAiUsage settles.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): schedule reconcile cron, drain backfill, bill errored-but-real AI spend

The "billed exactly once across crashes/deploys" guarantee had three holes:

W4a — reconcile cron unscheduled. /api/cron/reconcile-credits existed and was
HMAC-protected but absent from docker/cron/crontab, so nothing ever ran the
reconciliation. Added a signed GET every 10 minutes using the same cron-curl
pattern as the other ~16 jobs.

W4b — backfill never drained. backfillCredits() did a single LIMIT 200 pending
sweep + single LIMIT 200 orphan sweep; a backlog >200 silently left the rest.
It now loops until a pass returns fewer than BATCH from both sweeps (settled
rows drop out of the next query, so re-querying makes forward progress),
bounded by MAX_PASSES=50 as an unbounded-run backstop. GRACE_MS cutoff and the
isBillingEnabled() guard are unchanged; returns cumulative {retried, orphans}.

R1 — errored-but-real spend was dropped (deliberate billing-policy change).
Tokens consumed before a mid-stream error/abort are real provider cost, but
trackAIUsage only billed when success===true and the orphan sweep filtered
success=true, so an errored generation that produced tokens was logged with a
real cost and billed by neither path. Now:
  - trackAIUsage bills when aiUsageLogId && (success || totalTokens > 0). A
    token-less pre-generation failure still carries 0 tokens and is skipped;
    consumeCredits still settles a zero-charge call as 'skipped' (no $0 churn).
  - the orphan sweep reconciles success:false rows carrying cost > 0, and now
    filters gt(cost, 0) so no/zero-cost rows stay excluded.
The base PR intentionally left failed calls unbilled; the audit owner has
decided errored-but-real spend MUST be billed.

Tests: backfill drains a >200 backlog across passes, stops at the safety cap,
bills a success:false orphan with cost, and asserts the sweep no longer gates
on success; trackAIUsage bills an errored call with tokens but not a token-less
failure. @pagespace/lib typecheck clean; billing + ai-monitoring suites green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): fund prepaid balances from Stripe (invoice.paid refill + credit-pack top-up)

The pure routing/arithmetic in credit-core (classifyStripeEvent,
computeMonthlyRefill, applyTopup) had zero callers and there was no funding
shell, so a paid invoice or credit-pack purchase never became spendable credit.
The webhook only logged invoice.paid and ignored credit-pack checkouts.

Add credit-funding.ts — an imperative shell that:
  - invoice.paid -> resets the monthly bucket to the tier allowance, rolls the
    billing window forward from the invoice period, and records a monthly_grant
    ledger row keyed on the invoice id.
  - checkout.session.completed (mode=payment, kind=credit_pack) -> adds the pack
    to the never-expiring top-up bucket via applyTopup, recording a topup_purchase
    ledger row keyed on the session id.

Exactly-once: each funding ledger insert uses onConflictDoNothing against the
partial unique index credit_ledger_stripe_ref_unique (predicate restated as the
arbiter), and the balance mutation only runs when that insert actually inserted —
so a redelivered Stripe event credits the balance exactly once. Ledger insert and
balance write share one transaction. Funding never throws into the webhook: a
failure is logged and swallowed; Stripe retry / the reconcile cron re-delivers.
No-op when billing is disabled (tenant/onprem) and for tier_change events (tier
persistence stays in handleSubscriptionChange; the next invoice.paid refills at
the new allowance).

Wire applyStripeFunding into the webhook for invoice.paid and
checkout.session.completed, additively — existing logging and subscription-tier
behavior are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): enforce prepaid gate at AI entry points + log usage when metadata missing

Wire canConsumeAI() into every user-facing AI generation entry point so an
out-of-credits user is blocked (HTTP 402) before the model runs, instead of
the platform silently fronting the overage. Also fix the global-messages route
to always write an aiUsageLogs row (R4) so the orphan-sweep can recover/bill
calls where the provider returned no usage metadata.

Entry points gated (402 out_of_credits when !gate.allowed):
- api/ai/chat
- api/ai/global/[id]/messages
- api/v1/chat/completions
- api/ai/page-agents/consult
- api/pulse/generate (on-demand; cron path intentionally not gated)

R4: api/ai/global/[id]/messages now always calls AIMonitoring.trackUsage in
onFinish (0/undefined tokens are fine — $0 cost, but the log row exists).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): make credit-pack top-up funding race-safe; correct funding-failure docs

Self-review of the funding shell surfaced a lost-update on the top-up money path:
two concurrent first-time credit-pack purchases (distinct session ids, so their
ledger inserts don't serialize each other) would both SELECT ... FOR UPDATE on a
not-yet-existent balance row — locking nothing — both read 0, and the second
write would overwrite the first instead of adding to it, silently dropping a paid
top-up.

Fix: inside the funding transaction, ensure the balance row exists first
(INSERT ... ON CONFLICT DO NOTHING), then SELECT ... FOR UPDATE always locks a
real row, making the read-add-write atomic. applyTopup is still the source of the
new value; concurrent purchases now serialize on the row lock and both increments
apply. Add a first-time-buyer regression test alongside the existing add-to-
existing-balance test.

Also corrected the module/function docs: funding swallows its own errors (never
500s) and the webhook's coarse stripeEvents guard blocks same-event reprocessing,
so a failed funding event is not auto-recovered by Stripe retry. The previous
comment overclaimed retry/cron recovery; it now states failures are surfaced via
logs for operator follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): non-zero reserve floor, shortfall-as-debt, sub-cent remainder, gate-driven monthly reset

Closes four correctness gaps in the prepaid credit core:

R2a — Reserve floor default 0 -> 25¢, bounding the single in-flight call
that can overshoot zero (cost is only known post-stream). Env override kept.

R2b — Shortfall is no longer silently discarded. decrementAndSettle records
appliedCents (what actually left the balance) on the usage row and, when the
balance can't cover the charge, writes the uncovered remainder as a terminal
'adjustment' (debt) row in the same txn — visible, queryable by aiUsageLogId,
recoverable. Balances stay >= 0 (DB CHECK); debt lives in the ledger. The
usage-log unique index is scoped to entryType='usage' so the debt row can
share the call's aiUsageLogId.

R3 — Sub-cent costs no longer round to $0. Charges accrue in millicents into a
per-user pendingMillicents carry; each settle debits floor(pending/1000) whole
cents and banks the remainder. New pure core: chargeMillicents / accruePending
/ accrueCharge. No float ever reaches stored state.

W3-free — canConsumeAI now stamps a monthlyPeriod{Start,End} on lazy-init and,
when the window has expired, resets the monthly bucket to the tier allowance and
rolls the window forward — giving free/no-subscription users a monthly reset
without a cron. The reset UPDATE re-checks expiry in its WHERE so a racing
invoice.paid refill naturally wins.

Schema: + credit_balances.pendingMillicents, + credit_ledger.appliedCents,
+ credit_ledger.chargeMillicents; usage-log unique index scoped to 'usage'.
Migration 0144 generated (not hand-written).

Out of scope (tracked separately): per-user in-flight concurrency cap /
reservation, which requires threading a reservation id through routes+monitoring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): mark funding ledger rows settled so the backfill cron can't claw them back

Critical: monthly_grant and topup_purchase rows inherited creditLedger's default
consumeStatus 'pending'. backfillCredits() sweeps EVERY pending ledger row through
settlePendingLedgerRow() -> decrementAndSettle(), which SUBTRACTS abs(amountCents)
from the balance (it exists to settle unsettled *usage* charges). A funding row has
a positive amountCents, so after the 5-minute grace period the cron would reverse
every grant/top-up — clawing back exactly the credit funding just added.

Funding applies its balance change in the same transaction as the ledger insert, so
the row is already settled the moment it is written. Insert funding rows with
consumeStatus 'applied' so the pending sweep skips them. Add assertions to the
monthly-refill and top-up tests pinning consumeStatus 'applied'.

Reported by Codex review (P1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(billing): end-to-end credits-flow integration tests

Drive the real billing shells (applyStripeFunding, canConsumeAI,
consumeCredits, backfillCredits) wired together over one shared
in-memory DB, proving the prepaid money path fund→gate→consume→reconcile
works as a single system. Covers happy path, idempotency (aiUsageLogId +
stripeRef), crash recovery (pending settle, orphan sweep, success:false
billing, >BATCH multi-pass drain), monthly reset, sub-cent accrual,
shortfall/debt, and billing-disabled no-op.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(billing): align gate tests with #1473 (mock createAdminRestrictedResponse + isAdminOnlyProvider)

* fix(billing): await ask_agent metering so the sub-agent charge is durable

Addresses Codex P2: trackAIUsage now persists+debits inside its returned
promise, so the ask_agent call must await it (matching the chat/v1 handlers)
or the sub-agent usage log and credit debit can be dropped on a serverless
return.

* fix(billing): make funding failures retryable via Stripe redelivery

Codex P1 (re-raised): the webhook commits its stripeEvents idempotency marker
before processing, so a swallowed funding failure was lost forever — Stripe's
redelivery short-circuits as "already processed" and the backfill cron reconciles
only usage rows, not funding. A transient DB error during invoice.paid or a
credit-pack checkout could leave a paying customer permanently unfunded.

applyStripeFunding now logs and RE-THROWS genuine failures (non-actionable cases —
billing disabled, ignored events, unknown customer, missing ids — still return
quietly). The webhook wraps funding in fundOrLetStripeRetry: on a funding failure
it deletes the stripeEvents marker and rethrows, so the route returns 500 and
Stripe redelivers, reprocessing the event. Funding is idempotent on
creditLedger.stripeRef, so the balance is still credited exactly once.

Safe to reprocess: the handlers that run before funding on these events are
log-only / no-op (handleInvoicePaid only logs; handleCheckoutCompleted acts only
for mode 'subscription', whereas a credit-pack top-up is mode 'payment').

Tests: rethrow-on-failure assertion replaces the old swallow test; added a
non-actionable-cases test asserting no throw for unknown customer / ignored /
billing-disabled. 55 billing tests pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(lib): export ./billing/credit-funding so the web build can resolve it

The Stripe webhook route imports @pagespace/lib/billing/credit-funding, but the
new module was missing from package.json exports + typesVersions. tsconfig-path
typecheck resolved it locally, but the Next build resolves via package exports
and failed collecting /api/stripe/webhook. Verified: web build now exits 0.

* fix(billing): make the whole funding-relevant webhook path retryable, not just funding

Codex P1: the marker-cleanup only wrapped the funding call, but for invoice.paid
handleInvoicePaid runs first and does a DB user lookup. A transient failure THERE
threw before the cleanup, leaving the stripeEvents marker in place — the outer
catch returns 500, and Stripe's redelivery short-circuits at the marker conflict
and never reaches applyStripeFunding, so the paid monthly credit is still lost.

Replace fundOrLetStripeRetry(event) with withFundingRetry(eventId, run): it wraps
the WHOLE funding-relevant case body (pre-funding handler + applyStripeFunding) and
deletes the marker on ANY failure before rethrowing, so Stripe redelivers and the
whole path reruns. Safe to reprocess: handleInvoicePaid only logs;
handleCheckoutCompleted's sole throwable is an idempotent customer-link upsert (and
it acts only for mode 'subscription' — a credit-pack top-up is mode 'payment'; its
provisioning POST swallows its own errors); funding is idempotent on stripeRef.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): gate self-heals a NULL monthly period so top-up-first users get the free allowance

If a credit-pack purchase creates the first credit_balances row (topup funding
inserts { userId } only — monthly 0, monthlyPeriodEnd NULL) before the user's
first AI request, the gate previously skipped both the reset (period was NULL,
not expired) and lazy-init (a row existed), so the free monthly allowance was
never granted and never reset. The gate now resets on (period IS NULL OR
expired), stamping a window and granting the tier allowance. Race-safe: the
UPDATE re-checks the same predicate. +1 test; integration fake-DB engine gains
an 'or' operator.

* fix(billing): address CodeRabbit review (gate ordering, paid-user reset, migration CHECK, test rigor)

- chat route: run the prepaid gate BEFORE persisting the user message, so a 402
  no longer leaves an orphaned/duplicate prompt in chat history on retry.
- credit-gate: restrict gate-driven monthly reset to FREE/non-subscription users.
  Paid tiers refill authoritatively via invoice.paid; gate-resetting them would
  over-grant when a renewal invoice is late or retried. +tests (paid user blocked).
- migration 0144: add the credit_balances_pending_millicents_range CHECK so DBs
  upgraded through this migration enforce the 0<=pending<1000 carry invariant.
- global R4 test: prove trackUsage is AWAITED via a never-resolving deferred
  (a synchronous mock + 'was called' could not catch a fire-and-forget regression).
- integration test: derive the billing window from Date.now() so the funded
  period is always active (a hardcoded past window let the gate refill mid-test).
- consult test: drop the file-wide no-explicit-any disable; type the mock helpers.

* fix(billing): bound drain loop on no-progress, not just MAX_PASSES

Review hardening for the backfill drain loop:

- No-forward-progress break. `decrementAndSettle` (credit-consume.ts) leaves a
  ledger row 'pending' when the user has no balance row yet, so such rows never
  drop out of the pending sweep. The drain loop would therefore re-fetch and
  re-attempt the same unprocessable batch every pass up to MAX_PASSES (50× the
  work, every 10-min cron run, for a balance-less backlog). Now each pass
  fingerprints the fetched rows (order-independent); two identical consecutive
  passes mean nothing settled, so the loop stops instead of churning. A truly
  stuck full batch now ends in 2 passes, not 50. Partial progress still drains
  normally via the short-pass break.

- Hoisted the cap-exhaustion warning out of the loop body. It now fires exactly
  when the loop ran to MAX_PASSES (a real remaining backlog), not on a narrow
  last-iteration batch-size coincidence, and is now covered by tests.

Tests: +1 (stuck full batch stops early, no MAX_PASSES warning) and the cap test
now asserts the warning fires. @pagespace/lib typecheck clean; full lib suite
178 files / 4361 tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 2e993ec)

* fix(billing): normalize appliedCents to avoid storing -0 on sub-cent settles

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 34a6cd2)

* fix(billing): clamp month rollover in gate reset to avoid month-end overflow

CodeRabbit P2: setUTCMonth(+1) turns Jan 31 into Mar 3, making the 'monthly'
reset window longer than a month and delaying the next allowance refill for
users initialized/reset near month end. addOneMonth now clamps to the last
valid day of the target month (Jan 31 -> Feb 28/29). Exported + unit-tested
across mid-month, month-end, leap-year, and year-rollover cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 7e69435)

* fix(billing): address credits-remediation review findings (gate ordering, funding tier/customer, expired paid monthly, webhook test)

Five follow-ups from the /aidd:review of pu/credits-remediation:

- [P1] Mock applyStripeFunding in the webhook unit tests. The new funding call
  loaded the real billing module against a mock db that can't satisfy its query
  chain, turning every funding-relevant event into a 500. Mock the module and
  add wiring assertions (funding invoked once per event; 500 is retryable).

- [P1] Resolve a credit-pack buyer from trusted session metadata.userId before
  the customer link. A first-time payment-mode checkout doesn't link the Stripe
  customer to a user, so the customer lookup missed and the top-up was silently
  dropped. resolveTopupUser prefers metadata.userId, falls back to the customer.

- [P2] Run the global-assistant credit gate BEFORE persisting the user message,
  matching the page-chat route. A denied request no longer leaves an orphaned
  prompt that duplicates on top-up + retry.

- [P2] Exclude a paid user's expired monthly bucket from the gate decision. Once
  monthlyPeriodEnd has passed (renewal delayed), leftover monthly allowance no
  longer funds calls — only the never-expiring top-up does (blocked-until-renewal).

- [P2] Base the monthly refill on the tier derived from the PAID invoice line,
  not the stored users.subscriptionTier, so an invoice.paid that races ahead of
  the subscription webhook still grants the correct allowance. applyStripeFunding
  takes an optional { tier }; the webhook derives it via getTierFromPrice.

Tests: +4 lib billing (95 pass), +4 web route (webhook/global gate). lib + web
typecheck clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(marketing): move pricing/FAQ/terms copy to metered AI credits (#1495)

Rewrite every "N AI calls per day" surface in apps/marketing to the
prepaid metered AI-credits story: each tier includes a monthly $ AI-credit
allowance that meters usage, you can buy more anytime via top-up packs,
unused monthly credits reset each billing period, and model access still
differs by tier (free = standard models; paid = standard + Pro models).

Numbers are sourced from packages/lib billing/credit-pricing.ts via a new
single-source-of-truth module (apps/marketing/src/lib/credits.ts) so public
copy can't drift from what the app actually meters.

Surfaces updated: pricing page (cards + comparison table), FAQ (incl. new
"how credits work" entry + reworded out-of-credits answer), Terms (plan
list + usage-limits section), getting-started + features/ai docs, privacy,
schema.org offers, search index, and the BYOK blog post. Storage/file-size
copy left unchanged.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): AI unit-economics observability (margin queries + admin view) (#1494)

* feat(billing): AI unit-economics observability (margin queries + admin view)

Add margin aggregation queries joining aiUsageLogs with creditLedger usage
rows to surface real provider cost vs charged credits, gross margin %, and
uncovered debt per period, model/provider, and user.

- monitoring-queries.ts: computeMarginPct + getUnitEconomicsSummary,
  getMarginByPeriod, getMarginByModel, getTopSpendersByMargin,
  getOutstandingDebtByUser. Magnitudes via ABS(); debt summed from
  'adjustment' rows by ledger createdAt (no join) so retention purges
  can't under-report. Granularity is a bound param, not interpolated.
- GET /api/admin/unit-economics (withAdminAuth): JSON snapshot + CSV export.
- /admin/unit-economics admin view: summary cards, margin by model, top
  spenders, outstanding debt, margin-over-time; linked from /admin nav.
- Unit tests for margin logic, filters, and entryType scoping (14 tests).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): sum precise sub-cent fields in unit-economics aggregates

Summing per-row-rounded realCostCents/amountCents rounded high-volume
sub-cent traffic to $0 and reported bogus margin. Aggregate the precise
fields and round once: charged from SUM(chargeMillicents)/1000, real cost
from SUM(aiUsageLogs.cost)*100. appliedCents stays exact (whole-cent debit,
remainder banked in pendingMillicents). Debt keeps summing amountCents since
'adjustment' rows carry only the whole-cent shortfall.

Addresses CodeRabbit P2 on PR #1494.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(email): announce metered AI credits to users (template + broadcast script) (#1496)

* feat(email): announce metered AI credits to users (template + broadcast script)

Workstream F of the metered-AI-credits cutover: a one-time announcement
email telling users their AI usage is moving from daily call limits to a
monthly pool of prepaid AI credits (with buy-more top-ups).

- CreditsChangeEmail React Email template matching the existing template
  visual style; states the per-tier monthly allowance + top-up packs and
  reassures that documents/tasks/channels/collaboration are unaffected.
- credits-change-content helper derives all per-tier dollar figures straight
  from billing/credit-pricing (TIER_MONTHLY_ALLOWANCE_CENTS + CREDIT_PACKS)
  so the email can never quote a number the gate doesn't actually grant.
- render-email helper wraps @react-email/components render so repo-root
  scripts can produce email HTML without depending on it directly.
- send-credits-change-notifications.ts broadcast script mirrors
  send-tos-notifications.ts: queries all users with a valid email, sends via
  the shared rate-limited sendEmail, and is idempotent/resumable via a local
  JSONL ledger (re-runs skip already-sent recipients; failures retry).
  Supports --dry-run, --verified-only, --limit, --delay-ms, --log.

Verified end-to-end with --dry-run against a seeded DB: per-tier numbers
render correctly (free $5 / pro $15 / founder $50 / business $100), invalid
emails skip, ledger entries skip, and --verified-only/--limit behave. No
real emails sent. lint, typecheck, build, and lib tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(email): address Codex review on credits broadcast script

- Refuse a live (non --dry-run) send when the resolved app base URL points at
  localhost, so the broadcast can never email a broken CTA. Resolve the URL
  from NEXT_PUBLIC_APP_URL then WEB_APP_URL, preferring the first non-localhost
  value (handles a setup where only the server-side WEB_APP_URL is production).
- Make the idempotency ledger crash-safe: open + validate writability before
  the first send, fsync each record, and treat a ledger-write failure after a
  successful send as fatal — abort and name the unrecorded recipient so a
  re-run can never silently double-send it.
- Drop the unused emailVerified column from the user select.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): credit reservation holds + free in-flight cap; retire daily AI rate limits (#1497)

* fix(billing): credit reservation holds + free in-flight cap; retire daily AI rate limits

Workstream A — reservation/hold + free-tier in-flight cap:
- New credit_holds table (id, userId fk cascade, estCents, aiUsageLogId,
  createdAt, expiresAt; indexed on userId + expiresAt) via db:generate.
- credit-core (pure): reservationCents(), holdExpiresAt(), and evaluateGate
  extended to subtract reservedCents + estCostCents from spendable and to deny
  with a new too_many_in_flight reason when inFlightCount >= maxInFlight.
- credit-pricing: CREDIT_HOLD_ESTIMATE_CENTS (default = reserve floor),
  CREDIT_HOLD_TTL_SECONDS (900), MAX_FREE_INFLIGHT (2), all env-tunable.
- credit-gate canConsumeAI: authoritative decision now runs in one transaction
  that locks the balance row, sums & counts non-expired holds, denies the
  free-tier in-flight cap and out-of-credits, else inserts a hold and returns
  { allowed, holdId }. GateResult gains holdId.
- credit-consume: consumeCredits({…, holdId?}) releases the hold inside the
  settle transaction (and on the zero-charge path); new releaseHold() frees a
  reservation for token-less failures that never bill.
- credit-backfill reconcile: sweeps holds past expiresAt so a crashed stream's
  reservation can't permanently shrink spendable (BackfillResult.expiredHolds).
- holdId threaded gate -> route -> billing: AIUsageData/trackUsage ->
  consumeCredits, across all 5 AI routes (+ agent-communication-tools note).
  Shared credit-gate-response helper maps out_of_credits -> 402,
  too_many_in_flight -> 429.

Workstream B — retire daily AI rate limits:
- chat + global-assistant routes: removed the getCurrentUsage ->
  createRateLimitResponse (429) blocks and the incrementUsage/broadcastUsageEvent
  calls in onFinish. Model-tier gating (requiresProSubscription / admin-only
  providers) preserved. usage-service / rate-limit-cache / rate_limit_buckets /
  sweep-expired left intact — confirmed they back auth/login/integration limits.

Tests: extended billing unit + integration suites (hold accounting, in-flight
cap, hold release on settle, expiry sweep, full gate->consume hold lifecycle)
and added route 429 + helper coverage. lib 4426 + web routes green; typecheck,
lint, web build all pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): address Codex review — settle-time monthly expiry + release holds on early route exits

P2 (credit-gate.ts): use-it-or-lose-it now holds at SETTLE too. decrementAndSettle
excludes an expired monthly window (paid user past monthlyPeriodEnd) so allocateSpend
no longer silently draws the forfeited monthly allowance — it spends top-up only and
drops the stale monthly, matching the gate's exclusion.

P2 (global/[id]/messages + chat + consult + pulse routes): release the credit hold on
pre-generation early returns/throws. A holdHandedOff flag + finally frees the reservation
whenever the request exits after the gate but before the stream/billing takes ownership
(auth/permission/provider/save failures), instead of stranding it against the user's
balance + in-flight cap until the reconcile sweep. v1 unchanged (no explicit early return
after its gate; a throw falls through to the reconcile backstop like any crashed stream).

Also: export @pagespace/lib/billing/credit-consume (exports + typesVersions) so routes can
import releaseHold; add settle-time expired-monthly unit test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(db): regenerate credits migration onto master's 0144 (resolve migration-number collision)

Merging master brought in Canvas Publishing's 0144_flawless_ma_gnuci, colliding with
our hand-numbered 0144_big_valkyrie/0145_brave_exodus. Took master's 0144 as canonical
and regenerated a single 0145 capturing the credits schema delta (pendingMillicents,
appliedCents, chargeMillicents, credit_holds, usage-log index rescope) on top of it.
Re-added the pendingMillicents range CHECK by hand — drizzle-kit in this repo doesn't
emit/track CHECK constraints, so the regenerate would otherwise silently drop the
[0,1000) money-path invariant the original migration enforced.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): credit balance UI, buy-credits checkout, and out-of-credits UX (#1499)

* feat(billing): credit balance UI, buy-credits checkout, and out-of-credits UX

Workstream C + the in-app (apps/web) copy of Workstream D for the metered
AI-credits cutover. Gives users the surfaces that make the hard cutover humane:
see their balance, get correct 402/429 messages with a CTA, and buy more credits.

- Balance API: GET /api/credits → { monthly, topup, spendable, reserved } with a
  read-only getCreditBalance() that mirrors the gate's window semantics for display
  (free-tier lapsed → full allowance; paid lapsed → 0 monthly; spendable nets holds).
  SWR hook useCreditBalance with live socket updates.
- Live updates: replace the retired daily-quota usage:updated socket event with
  credits:updated (broadcastCreditsEvent + emitCreditsUpdated), emitted after a call
  settles (the two interactive AI routes) and after funding (the Stripe webhook).
- Widget: replace UsageCounter with a CreditBalance header widget (remaining + low
  warning + Buy credits) and a CreditBalanceCard on settings/billing. settings/plan
  + settings/billing now show credits, not aiCalls/day.
- Buy-credits checkout: POST /api/stripe/create-credit-topup mirrors create-subscription
  (mode:'payment', inline price_data from CREDIT_PACKS, metadata.kind='credit_pack'),
  so the existing webhook funds the top-up bucket. BuyCreditsButton in settings + in
  the out-of-credits chat error states.
- Error UX: classifyAIError distinguishes out_of_credits (402) and too_many_in_flight
  (429) with distinct copy; SidebarChatTab + ChatInputArea show a Buy-credits CTA.
- Copy: plans.ts limits move from aiCalls/pro/day to a monthly credit allowance +
  proModels capability, sourced from credit-pricing via a new web credits.ts helper
  (mirrors apps/marketing/src/lib/credits.ts). Retire the orphaned /api/subscriptions/usage.

Tests: balance logic, GET /api/credits, the top-up route, error classification, and
the renamed socket event. Lint + web typecheck + web build green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): hide BuyCreditsButton on iOS / billing-disabled (Codex P2)

The out-of-credits error CTAs (ChatInputArea, SidebarChatTab) rendered
BuyCreditsButton unconditionally, exposing a Stripe checkout on iOS Capacitor
builds where billing UI must be hidden for App Store compliance. Make
BuyCreditsButton self-hide via useBillingVisibility so every call site —
including the error states — is compliant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ai): dedupe users import in v1 completions after merging master

master #1500 (server-side tool execution) and our credit gate both added
`import { users }` to the completions route; the clean text-merge left a
duplicate identifier (TS2300). Removed the redundant import.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): address CodeRabbit review on metered-credits PR

- v1 completions: release the credit hold on stream failure (was leaking the
  reservation until TTL/reconcile, leaving users artificially short).
- chat route: create the conversation row AFTER the credit gate so a denied
  first prompt leaves no orphaned conversation.
- monitoring-queries: anchor the unit-economics window on creditLedger.createdAt
  and LEFT JOIN aiUsageLogs (was inner-join + usage-log timestamp, which dropped
  charged credits/margin once a usage log was retention-purged); bucket
  purged-log rows under 'unknown' model/provider. Test updated to match.
- admin CSV export: neutralize spreadsheet formula injection (=,+,-,@) in
  attacker-controlled name/email cells.
- error classifier: tighten to exact codes/phrases so "context window limit
  exceeded" / generic "ai credits" no longer misroute to rate-limit/buy-credits.
- admin unit-economics page: render period buckets from the server string
  (no Date reparsing) to avoid timezone-shifted / mislabeled month buckets.
- credits route: validate subscriptionTier at runtime instead of casting.
- create-credit-topup: reject malformed/non-object JSON with 400, not 500.
- AiUsageMonitor: only compare ids this monitor is scoped to (page-agent mode
  has no conversationId, so it was dropping every credits:updated event).
- send-credits-change-notifications: count ATTEMPTS against --limit so a
  provider outage can't blow past a canary cap.

Not changed: CodeRabbit's "decouple apps/marketing from @pagespace/lib" — master
already couples it (contact route imports @pagespace/lib/security), so the
standalone premise is stale; the import is the no-drift source of truth. Replied
on-thread.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(billing): CREDITS_ENFORCEMENT_ENABLED kill-switch (dark-launch the gate)

The metered-credits cutover otherwise hard-enforces (live 402/429) the moment it
deploys, on placeholder allowances. Add an env flag so the gate can be dark-launched:
deploy the code in observe-only mode (meter + record real cost/charged credits for the
unit-economics view) and flip blocking on deliberately once the numbers are validated.

- credit-pricing.ts: envBool + isCreditsEnforcementEnabled() (default FALSE), read at
  call time so it toggles via env+redeploy and is settable per-test.
- credit-gate.ts: the gate still does ALL bookkeeping (lazy-init, monthly reset, balance
  read, hold on the allow path); when enforcement is OFF it only overrides a would-be
  denial (out_of_credits / too_many_in_flight) to allowed:'enforcement_disabled'. A
  credit-having user is unchanged (normal allow + hold). consumeCredits is untouched, so
  metering/observability run regardless.
- credit-core.ts: add 'enforcement_disabled' to GateReason (an allowed reason;
  credit-gate-response only maps deny reasons, so no HTTP change).
- tests: the two suites that exercise real enforcement set CREDITS_ENFORCEMENT_ENABLED=true;
  new dark-launch cases assert denials are suppressed while bookkeeping still runs.

To enforce in production: set CREDITS_ENFORCEMENT_ENABLED=true and redeploy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(db): rebase credits migration onto master's 0145 (resolve 0145 collision)

Merging master brought Agent Code Execution's 0145_flawless_living_mummy (#1487),
colliding with our regenerated 0145. Took master's 0145 as canonical and regenerated
the credits delta as 0146_furry_abomination via db:generate. Verified mechanically:
0146.prevId == master 0145.id, 0145.prevId == master 0144.id, journal idx + when
timestamps strictly increasing — chain points at the correct parent. Re-added the
pendingMillicents range CHECK (drizzle-kit doesn't emit CHECK constraints).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(billing): exempt numeric cells from CSV spreadsheet-injection guard

sanitizeSpreadsheetCell ran on every cell, so legit negative exports (marginUSD
"-0.37", marginPct "-12.50") got quote-prefixed and landed as text in Excel/Sheets.
Exempt plain numbers (/^-?\d+(\.\d+)?$/); keep quoting only non-numeric text that
starts with a formula trigger (=,+,-,@), i.e. the name/email columns.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

This branch was previously deployed

1 inactive deployment
Preview — 8b0efa7a Deployed Jun 2, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant