Skip to content

kafka connect: wait through connect-gate cold start instead of failing - #2666

Draft
simonlord wants to merge 2 commits into
masterfrom
sl/kafka-connect-starting-503
Draft

simonlord wants to merge 2 commits into
masterfrom
sl/kafka-connect-starting-503

Conversation

@simonlord

@simonlord simonlord commented Oct 6, 2026 •

Copy link
Copy Markdown

What and why

Redpanda Cloud is adding connect-gate (redpanda-data/cloudv2#30703), a proxy in front of Kafka Connect that lets the workers scale to zero when idle. While the workers boot, the gate answers mutating calls with 503, a Retry-After header and a body like:

{"error_code":503,"reason":"kafka_connect_starting","phase":"pulling-image","retry_after_seconds":60,"estimated_wait_seconds":180,"message":"Kafka Connect is starting (pulling the Kafka Connect image). Retry in about 3m0s."}

Today Console turns that into a terminal red error (and the validate endpoint even returns HTTP 200 for it), so a user creating the first connector after a quiet period sees "Failed to validate Connector config: … (503)" with no hint that waiting a few minutes fixes it.

Backend

  • A transport wrapper on the connect-client captures the gate's body and Retry-After per request (the client itself keeps only error_code and message), carried as connect.StartingError.
  • REST handlers for create / validate / update / delete / pause / resume / restart / restart-task answer 503 with Retry-After and reason, phase, retryAfterSeconds, estimatedWaitSeconds in the body. Other errors are unchanged.
  • Connect-RPC KafkaConnectService (v1, v1alpha2) maps 503 to CodeUnavailable with RetryInfo and ErrorInfo metadata instead of CodeInternal.
  • Validate now propagates the upstream status like the other calls instead of hardcoding 200.

Frontend

  • retryWhileKafkaConnectStarting retries a call while it fails with reason: kafka_connect_starting, honouring the retry hint (15 min budget), with a per-second countdown callback.
  • Create wizard: validate and create wait through the cold start with a "Kafka Connect is starting… retrying in Ns" warning (Review step and the creating dialog). The Properties step's initial validation does the same; validate-on-change skips quietly while cold.
  • Connector action confirm dialog (pause/resume/restart/delete/update): same waiting state, retry on its own.
  • Fixes a pre-existing bug: a validate failure in the Review step threw past the wizard and left it on the loading skeleton; it now renders "Validation attempt failed".

Examples

REST, while the workers boot:

PUT /api/kafka-connect/clusters/redpanda/connector-plugins/org.apache.kafka.connect.mirror.MirrorSourceConnector/config/validate
HTTP/1.1 503 Service Unavailable
Retry-After: 60
{"statusCode":503,"message":"Kafka Connect is starting (pulling the Kafka Connect image). Retry in about 3m0s.","reason":"kafka_connect_starting","phase":"pulling-image","retryAfterSeconds":60,"estimatedWaitSeconds":180}

Connect-RPC: unavailable with RetryInfo{retry_delay: 60s} and ErrorInfo{reason: REASON_KAFKA_CONNECT_API_ERROR, metadata: {reason: kafka_connect_starting, phase: pulling-image, retry_after_seconds: 60}}.

UI recording: to be added from an integration cluster running connect-gate (cloudv2 integration pack 26.2.20261006181804 bundles this branch as console-unstable:sl-kafka-connect-starting-503-8480a75).

Tests: go test ./pkg/connect/... (transport capture, ordinary 503 untouched, header/default fallbacks), src/utils/kafka-connect-starting.test.ts.

🤖 Generated with Claude Code

simonlord and others added 2 commits October 6, 2026 18:43
Kafka Connect in Redpanda Cloud may sit behind connect-gate, which scales
the workers to zero when idle and answers mutating calls with a 503 while
they boot. The connect-client reduces that to a code and message, so the
retry hints were lost and the REST layer returned a plain failure; validate
even reported HTTP 200.

Capture the gate's response in a transport wrapper, scoped to the request
through its context, and carry it as a StartingError. The REST handlers
then send 503 with Retry-After plus reason/phase/retryAfterSeconds in the
body, and the Connect-RPC services return CodeUnavailable with RetryInfo
and ErrorInfo metadata instead of CodeInternal. Validate now uses the
upstream status code like the other calls.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
When Kafka Connect has been scaled to zero, its first mutating call gets
a 503 with reason kafka_connect_starting and a retry hint. The create
wizard, the properties step and the connector action dialog now show a
"Kafka Connect is starting… retrying in Ns" state and retry on their own
until the workers answer, instead of a terminal error.

Also catches validation failures in the Review step: the throw there
escaped the wizard and left the step on its loading skeleton.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

🚨 Registry drift detected

App: frontend · Scope: diff vs origin/master · Files: 8

Count
⚠️ Outdated registry components 2
🛠 Locally-modified components 0
❓ Unknown to registry 0
🎨 Off-token palette colours 0
🔢 Ad-hoc utility classes 0
Components needing attention
Status Component Uses Detail
⚠️ outdated code-block 1× installed 3.5.0 → latest 3.6.0
⚠️ outdated stat 1× installed 3.4.1 → latest 3.6.0

Refresh command:

bunx shadcn@latest add @redpanda/code-block @redpanda/stat --overwrite

Generated by lookout audit-changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant