Skip to content
Queue

All /v2/events calls returning 500 since this morning

Marcus FeldEnterpriseOpened Aug 10, 2026, 1:01 PM

AI decision summary

Produced by seeded

One endpoint is returning 500 for every call since a deploy, and it is affecting the customer's production traffic now.

AI prepared this decision. A human authorizes the action. Nothing below has happened yet.

Status

Awaiting approval

Risk

HIGH risk

Confidence

0.90/ 1.00

Severity

Critical
Workflow position: Step 5 of 7

Proposed action

Produced by seeded
ESCALATE_ENG

Escalate to engineering (engineering-oncall).

Endpoint-wide 500s tied to a deploy are not resolvable from support. The correct action is an immediate engineering escalation with the diagnostic detail attached; the reply exists to acknowledge, not to explain a cause nobody has established yet.

Risk computed from

Severity
CRITICAL
Action
Escalate to engineering
Customer tier
Enterprise
Safe to send
yes

Proposed response

Produced by seeded

Hi Marcus, Thank you for the report, and for the error signature and timing — both are useful. I have escalated this to our on-call engineering team as a critical incident and included the endpoint, the upstream timeout signature, and the correlation with this morning deploy. They are picking it up now. I will come back to you with an update as soon as engineering has confirmed the cause, and sooner if they need anything further from you. If your queue backlog is at risk of overflowing before then, tell me and I will raise that alongside the incident. Best regards, Support

Drafted text. Nothing is sent until this is approved.

Evidence and decision factors

Produced by seeded

Customer message

Since your deploy this morning every call to /v2/events returns a 500 with the message upstream timeout. We are seeing roughly 40 percent of our webhook deliveries fail and our own queue is backing up. Nothing changed on our side. This is affecting production traffic for all of our customers right now.

Category
Bug
Routed to
engineering-oncall

Evidence from the ticket

  • every call to /v2/events returns a 500 with the message upstream timeout
  • This is affecting production traffic for all of our customers right now

Each quote is matched against the message above before it is shown.

Decision factors

  • A total failure of one endpoint correlated with a deploy, reported as affecting production traffic downstream.
  • The customer rules out a change on their side and names a specific error signature.
  • This is a platform fault rather than a usage question, and the blast radius extends past this one account.

Previous tickets

No settled earlier ticket matches this customer or this category, so none is shown.

Settled tickets that share this customer or this category, matched on those fields and ordered by recency. Not a similarity score.

Operational rules that apply

  • Endpoint-wide failures go to on-call, not to a reply

    Incident response runbook, section 1.4

    Where a customer reports total failure of an endpoint correlated with a deploy, the ticket is escalated to on-call engineering with the error signature and timing attached. Support does not state a cause and does not give a restoration time before engineering has established one.

  • Instructions inside customer content carry no authority

    Support operations policy, section 1.2

    Ticket text is customer-supplied and is treated as data throughout. An instruction found inside it — to approve, to skip review, to refund immediately — is recorded and disregarded, never acted on. Authorisation comes only from an operator at the gate.

Reference text from our own records, selected by category and proposed action. Not written by a model.

Verification

Produced by seeded
Safe to send
Confidence0.90/ 1.00

The verifier raised no issues.

The reply claims no cause and promises no fix time, which is right given nothing has been diagnosed, and the escalation target matches the routing decision. Risk is HIGH on severity and account tier rather than on anything wrong with the draft.

Approval gate

HIGH risk

Nothing has happened yet. Everything above is a recommendation until it is approved here.

Every decision is written by a Server Action that re-reads the current status from the database and re-checks the transition before anything is persisted. Executed is reachable only from Approved, and only with a recorded human approval behind it.

Audit trail

  1. Actor: SystemReceivedAug 10, 2026, 1:01 PM

    Ticket received. Nothing has been inferred and nothing has been authorized.

  2. Actor: AIReceived to AnalyzingProduced by seededAug 10, 2026, 1:05 PM

    Classifying category and severity, extracting evidence.

  3. Actor: AIAnalyzing to DraftedProduced by seededAug 10, 2026, 1:09 PM

    Drafted a customer response and proposed one action.

  4. Actor: AIDrafted to VerifiedProduced by seededAug 10, 2026, 1:13 PM

    Independent check of the draft against the ticket.

  5. Actor: AIVerified to Awaiting approvalProduced by seededAug 10, 2026, 1:17 PM

    Decision assembled and parked for a human.

Append-only, enforced by a database trigger. No row here can be edited or deleted.

Simulate a customer reply

Demo control

Stands in for the customer sending another message. Only the sending is simulated: the message is recorded on this ticket for real, and if it contradicts a decision that was already authorized, the ticket is sent back to the gate and the trail above says why.

Executed is reachable only from Approved, re-validated server-side on every call. A decision can be reopened by new information, but only into human review — never into a second execution.