Outcome
Make the first deployed API observable enough to diagnose startup, request and deployment failures without introducing the later distributed tracing scope.
Target iteration
Iteration 3, after automated App Service deployment works.
Acceptance criteria
Dependencies
Out of scope
- Cross-service tracing
- Service Bus telemetry
- Operational dashboards and production alerts
Completion and verification — 1 October 2026
Implemented in PR #72 and deployed at commit 2bb7cb5d2777c753df535ee96cb900623b132a61.
Live evidence
Verified through Application Insights and App Service Log stream:
- Startup:
TradingEngine.Api.Started events were recorded with the deployed version and commit, including the startup at 10:05:57 UTC.
- Requests and build metadata:
GET /health returned 200 and appeared with application.version, application.commit, and deployment.environment = Production. This is the ASP.NET Core environment name for the development App Service.
- Deliberate handled failure and SQL correlation: requesting the synthetic missing instrument
11111111-1111-1111-1111-111111111111 at approximately 10:36:36 UTC returned 404. Its transaction included a successful SQL dependency taking 235 ms, sharing operation ID c26861150519a26c54ff50ca6ed92e39. This verifies the story's handled-failure criterion using a missing-record 404; the runbook's alternative POST/400 example was not exercised.
- Exception diagnostics: the earlier 500 at approximately 10:21:59 UTC was correlated with a SqlClient connection timeout and useful exception type, messages, stack frames and source locations. Operation ID:
3e4ec69d2af25c347ade3f382f7ef635.
- Application output: App Service Log stream connected and displayed application logs, including Azure Identity token-cache diagnostics.
- Synthetic query/header check: a
GET /health carrying harmless synthetic values in a query parameter and X-Database-Probe-Key returned 200. Its request was ingested at 11:01:47 UTC with operation ID 446a3b619597432eb064a0a535b25a64; the stored URL had no query string. Searching requests, dependencies, traces, exceptions and customEvents for marker te36-sensitive-37b815f8f5bf returned no matches after the request was confirmed present.
Implementation evidence
The reviewed implementation provisions workspace-based Application Insights through Bicep, supplies its connection string through secure module wiring without exposing it in deployment outputs, and enables App Service diagnostics. Documentation covers application/deployment log locations, startup and request investigation, KQL and verification.
Development telemetry uses 30-day retention and a 0.1 GB daily ingestion cap; documentation explains that the cap can overshoot and is not a budget guarantee. Automated tests cover telemetry enrichment, sensitive-value handling, retained exception diagnostics and the Azure export boundary. SQL statement capture is suppressed, and public tests use synthetic data.
Separate follow-up
Review bounded SQL connection retry handling and serverless resume behaviour. The initial connection timeout was followed by a successful database call on retry. Auto-resume is a plausible explanation, but has not been confirmed from the database Activity log. This is a resilience follow-up and does not block completion of this observability story.
Outcome
Make the first deployed API observable enough to diagnose startup, request and deployment failures without introducing the later distributed tracing scope.
Target iteration
Iteration 3, after automated App Service deployment works.
Acceptance criteria
/healthrequests and one deliberate handled failure are visibleDependencies
Out of scope
Completion and verification — 1 October 2026
Implemented in PR #72 and deployed at commit
2bb7cb5d2777c753df535ee96cb900623b132a61.Live evidence
Verified through Application Insights and App Service Log stream:
TradingEngine.Api.Startedevents were recorded with the deployed version and commit, including the startup at 10:05:57 UTC.GET /healthreturned 200 and appeared withapplication.version,application.commit, anddeployment.environment = Production. This is the ASP.NET Core environment name for the development App Service.11111111-1111-1111-1111-111111111111at approximately 10:36:36 UTC returned 404. Its transaction included a successful SQL dependency taking 235 ms, sharing operation IDc26861150519a26c54ff50ca6ed92e39. This verifies the story's handled-failure criterion using a missing-record 404; the runbook's alternative POST/400 example was not exercised.3e4ec69d2af25c347ade3f382f7ef635.GET /healthcarrying harmless synthetic values in a query parameter andX-Database-Probe-Keyreturned 200. Its request was ingested at 11:01:47 UTC with operation ID446a3b619597432eb064a0a535b25a64; the stored URL had no query string. Searching requests, dependencies, traces, exceptions and customEvents for markerte36-sensitive-37b815f8f5bfreturned no matches after the request was confirmed present.Implementation evidence
The reviewed implementation provisions workspace-based Application Insights through Bicep, supplies its connection string through secure module wiring without exposing it in deployment outputs, and enables App Service diagnostics. Documentation covers application/deployment log locations, startup and request investigation, KQL and verification.
Development telemetry uses 30-day retention and a 0.1 GB daily ingestion cap; documentation explains that the cap can overshoot and is not a budget guarantee. Automated tests cover telemetry enrichment, sensitive-value handling, retained exception diagnostics and the Azure export boundary. SQL statement capture is suppressed, and public tests use synthetic data.
Separate follow-up
Review bounded SQL connection retry handling and serverless resume behaviour. The initial connection timeout was followed by a successful database call on retry. Auto-resume is a plausible explanation, but has not been confirmed from the database Activity log. This is a resilience follow-up and does not block completion of this observability story.