Skip to content

Customer-aware AI operations“The new version looks better on average. Is it safe for every customer?”

Ship AI changes without breaking a customer promise.

Synvolv connects a proposed change to the customers it can affect, replays it against production behaviour, applies customer policies, contract rules, entitlements and economics, then gives your team the decision: who moves, who stays, who needs review, why, and the rollout to approve

Connect with what you already run

  • GitHub
  • GitLab
  • OpenTelemetry
  • Anthropic
  • Stripe
  • Snowflake
  • Postgres
  • Braintrust

support-agent: move default model from GPT-5 to Claude Sonnet 5 #1816

Openrelease-eng wants to merge 3 commits into main from support-agent-v17

Some checks were not successful1 failing, 2 successful checks
  • ci / build — Successful in 2m 14sDetails
  • ci / test — Successful in 4m 02sDetails
  • Synvolv · customer-impact review — Do not release globally · 41 accounts affectedDetails

Synvolv customer-impact review

Change detected: support-agent v16 → v17 · 41 customer accounts · 12,420 production scenarios replayed

Works

Basic · English FAQPro · English FAQ

Regresses

Spanish refund workflow

Not permitted

Meridian HealthProvider policy

Not entitled

6 Enterprise accountsPremium-model requirement

Needs more evidence

2 accountsInsufficient replay coverage

Do not release globally.

Release to 27 eligible customer accounts

Customer-impact review

Your PR should know which customers it changes.

A pull request can change far more than code. It can change a model, a prompt, retrieval, tool permissions, retry logic, agent steps, routing, fallback behavior, or the product logic around the agent. Synvolv detects the relevant change, maps it to the production workloads and customers that can be affected, and replays representative production scenarios in shadow.

support-agent: move default model from GPT-5 to Claude Sonnet 5 #1816

Openrelease-eng wants to merge 3 commits into main from support-agent-v17

Some checks were not successful1 failing, 3 successful checks
  • ci / build — Successful in 2m 14sDetails
  • ci / test — Successful in 4m 02sDetails
  • lint — Successful in 41sDetails
  • Synvolv · customer-impact review — Do not release globally · 41 accounts affectedDetails

Synvolv customer-impact review

Change detected: support-agent v17 · GPT-5 → Claude Sonnet 5 · Affected: 41 customer accounts · Replay: 12,420 production scenarios

Works

Basic · English FAQPro · English FAQ

Regresses

Spanish refund workflow

Not permitted

Meridian HealthProvider policy

Not entitled

6 Enterprise accountsPremium-model requirement

Needs more evidence

2 accountsInsufficient replay coverage

Do not release globally.

Release to 27 eligible customer accounts

Not only

Quality +0.8%.

One number, standing in for every customer at once.

But

  • 01Which customers improved?
  • 02Which regressed?
  • 03Who is not entitled to the new behavior?
  • 04Which customer rules block it?
  • 05What happens to cost and latency?
  • 06Which accounts need more evidence?
  • 07What scope should actually be released?

Synvolv answers all seven, per customer, before the change ships.

Shadow mode

See what Synvolv would decide before it can touch anything.

Shadow mode is read-only. Synvolv reads a copy of what production already does, replays your next change against it, and shows which customers would move, which would stay, and why. Nothing changes for your users until you approve it.

Shadow reads from

  • GitHubGitLabGitHubthe pull request: what is about to change
  • OpenTelemetryProduction telemetrywhat ran, for which customer
  • StripeHubSpotCustomer identityplans, entitlements, contracts
  • BraintrustEvaluation / outcome signalshow it scored, how it resolved
  • Configuration historyprompts, models, tools and rules over time

Connect what you already run. Availability is confirmed per connector during setup.

Shadow

Synvolvread-only copy

Reads a copy of what production already does. Writes nothing back.

  • Review changes.
  • Replay historical scenarios.
  • See affected customers.

No write path to production

Would have decided

  • AcmeMoveinside threshold
  • Meridian HealthStayprovider policy
  • VelaReviewnot enough replay coverage
  • Read-only
  • No production write path required

Start in shadow. Keep your existing stack. No production action happens without approval.

The decision

One change. A different answer for every customer.

The same update to your AI can be right for one customer and wrong for the next. Synvolv checks each customer before anything ships, and tells you what to do with each one.

The change

Support agent v17

GPT-5 → Claude Sonnet 5

  1. AcmeShip itBetter and cheaper on their traffic
  2. Meridian HealthHoldTheir BAA covers Azure OpenAI only
  3. VelaCheck firstNot enough evidence it works for them
  4. MonarchShip with a capAlready past their included credits

Three checks decide it, for every customer: Does it work? Is it permitted? Does the account stay healthy?

Aggregate vs accounts

Your customers aren’t averages.

  • A model can win the eval.
  • A prompt can improve the overall score.
  • A retrieval change can reduce latency.
  • A new workflow can lower total cost.
  • The aggregate dashboard can stay green.

And the global rollout can still be wrong.

Overall

+0.8pt
Quality
−28%
Cost
−14%
Latency

Looks like an easy release.

  1. 01AcmeInside thresholdCost −31%Move
  2. 02Meridian HealthProvider policyNew provider prohibited by current policyStay
  3. 03Enterprise cohortPlan entitlementPremium model included in planStay
  4. 04Spanish supportRegressesResolution −8.2 ptsStay
  5. 05NorthstarStableCost −22%Move
2 move · 3 stay

The average was right.The global decision was wrong.

Synvolv makes the decision at the level where the consequence actually happens: the customer.

When you find out

Did your customers like the change?

You should not need an angry account, a support escalation, or the month-end margin report to find out.

A production change can look healthy overall while:

  • onetool starts looping
  • onehigh-usage account becomes expensive to serve
  • onecustomer should have stayed on the previous version

Move those discoveries before the rollout.

Synvolv connects the production evidence and customer context before your customers become the feedback loop.

How the call gets made

This is how fast-moving teams make the call today.

GitHub knows what changed. Your traces know what ran. Your eval platform knows how it scored. Customer Success remembers which customer is different, Sales knows what Enterprise bought, Legal knows what the agreement says, and Finance has the margin spreadsheet. And the person expected to say yes or no is usually an engineer or product lead.

Today the answer is assembled by hand

  1. Priya Raman10:42 AM

    The new version looks better on average. Is it safe for every customer?

  2. Marcus Lee10:53 AM

    Evals say +0.8 overall. Does every plan include the new behavior? @dana what did Enterprise buy?

  3. Dana Whitfield11:02 AM

    Enterprise is entitled to GPT-5. Not sure whether v17 changes that.

  4. Tomás Herrera11:09 AM

    Meridian is different. I think their BAA only covers Azure OpenAI. @legal has the contract.

  5. Aisha Bello11:15 AM

    The margin spreadsheet is in the shared drive, last updated Friday.

  6. Priya Raman11:20 AM

    Six tabs open (#1816, traces, the eval run). Building a cohort now. I’ll write a rollout rule so we can still ship this week.

The startup still needs to ship this week. So the team makes the best call it can. And if something breaks, everyone reconstructs the decision again in reverse.

With Synvolv the same question, answered

The new version looks better on average. Is it safe for every customer?

Synvolv answered from your connected systems

Read 4 sources · Evals · Entitlements · Contracts · Traces

Not yet.

47 customer accounts affected

Evidence 5

31 ready
Quality remains inside threshold · Cost / successful resolution −18%
Evals
8 stay
Enterprise plan is entitled to GPT-5
Entitlements
4 stay
BAA does not cover Anthropic yet
Contracts
2 regress
Spanish refund resolution −7.1 pts
Evals
2 review
Not enough replay coverage
Traces

Release scope ready

31 customer accounts · Basic + Pro · English FAQ

Start with 5% canary · Rollback per customer · Approval required

The context stops living in eight places. They meet at the decision your team is trying to make.

Where the answers come from

The context Synvolv holds for your company. You manage it.

Policies, contracts, plans, entitlements, budgets and data rules live in one knowledge base per company: synced from the systems that own them, or entered by your team. Every decision above points back to the entry it used, and you can change any entry.

  • Synced from your systems, or entered by hand
  • Every entry shows where it came from and who owns it
  • One record per customer, one per company

Meridian Health

Knowledge base

9 entries4 synced · 5 set by your teamSynced 2 min ago

Provider policyAzure OpenAI only, per the signed BAAContracts syncLegal3d
Scope
All workflows · Meridian Health
Effective
Since contract v3
Verified
Legal · contract v3, clause 4.2 (BAA subprocessors)
Cited by
Release decision · v17 Claude Sonnet 5 → Hold

History

  • 3dLegal confirmed clause 4.2
  • 5dSynced from contracts
  • 4wEntry created
PlanEnterprise · GPT-5 and GPT-5 miniBilling syncAisha Bello12m
Quality thresholdResolution may not fall more than 2 ptsSet by handDana Whitfield5d
Daily agent budget$1,200 · then GPT-5 miniSet by handAisha Bello1d
Data rulesCustomer content: opt-in · metadata by defaultSet by handPriya Raman2w

Synvolv makes that customer context part of the system.

The same four facts the thread went looking for. Where each one lives today, and where it lives once it is part of the system.

The factTodayWith Synvolv
PoliciesA clause in a contract PDF. Someone asks Legal to go and find it.Policies do not stay trapped in PDFs.The clause is an entry the decision reads, with a source and an owner.
EntitlementsA field in billing. Someone half-remembers what the plan included.Entitlements do not stay trapped in billing.What each plan includes is read from billing and applied per customer.
Customer exceptionsOne person's memory. If they are out this week, it is gone.Customer exceptions do not stay in someone's memory.They are written down, with an owner and a date, and checked on every change.
Production evidenceAn aggregate dashboard and six open tabs.Production evidence does not stop at an aggregate dashboard.The change is replayed per customer, so the evidence is theirs, not an average.

Today: eight places, one thread, a best call.With Synvolv: one record, one decision, with the evidence.

The product, end to end

What working with Synvolv looks like.

Each flow starts where your team already works — a pull request, a Slack channel, the Synvolv app — and shows what Synvolv does next.

Customer context

Your customer agreement changed. Does production know?

Those facts can determine which model is allowed, which provider can be used, which tools the agent may call, what quality must be maintained, how much usage is included, which release the customer can receive, and what happens when a limit is crossed.

  1. ContractSales

    Master services agreement · v3

    • 4.2 Subprocessors — model providers must be covered by the BAA
    • 5.1 Service level — Spanish resolution ≥ 92%
    • 7.3 Data handling — no customer content used for training

    A contract is not only something Sales closes.

  2. PlanBilling

    Enterprise · Support agent

    • Model entitlement — GPT-5, GPT-5 mini
    • Maximum cost / resolution — $0.60
    • Included usage — 250,000 resolutions a year

    An entitlement is not only something Billing stores.

  3. PolicySecurity

    Customer policy · Meridian Health

    • Providers — Azure OpenAI only, per the signed BAA
    • Tools — refunds need human approval
    • Fallback — GPT-5 mini when a limit is crossed

    A customer policy is not only something Security signs off.

Meridian Health

Customer record

Plan
Enterprise
Workload
Support agent
Provider policy
Azure OpenAI only · per BAA
Model entitlement
GPT-5 · GPT-5 mini
Spanish resolution floor
92%
Maximum cost / resolution
$0.60
Current release
v16 · GPT-5
Incoming release
v17 · Claude Sonnet 5

Rules in force 7

  1. which model is allowedGPT-5 · GPT-5 miniPlan
  2. which provider can be usedAzure OpenAI onlyContract
  3. which tools the agent may callRefunds need approvalPolicy
  4. what quality must be maintainedSpanish resolution ≥ 92%Contract
  5. how much usage is included250,000 resolutions a yearPlan
  6. which release the customer can receivev16 · v17 blockedPolicy
  7. what happens when a limit is crossedFall back to GPT-5 miniPolicy

The gap

What the team remembers

Nobody from that room is on the pull request.

  • Sales closed the renewal
  • Legal wrote clause 4.2
  • Security signed the BAA

What Synvolv holds

v17 held for Meridian Health

  • Clause 4.2 — Azure OpenAI only, per the BAA

Six months later, someone changes the product.Those facts still have to matter.

Synvolv keeps them connected to the customers and workloads they apply to.

  • approved customer policies
  • contract terms
  • plan rules
  • entitlements
  • quality thresholds
  • economic limits

The engineer does not have to remember this.Synvolv applies it to the decision.

Versioned rules

And when the customer rule changes, the decision changes too.

Synvolv is not a permanent “policy says no” gate. Imagine Meridian adds the new provider to its BAA and approves it.

Policy update

Provider policy · Meridian Health

v3 → v4Approved

The change

  • Azure OpenAI only, per the BAA
  • Azure OpenAI · Anthropic on AWS Bedrock
Source
BAA amendment 2 · countersigned
Requested by
Meridian Health · Legal
Approved by
Tomás Herrera · Security review
Condition
Spanish resolution must remain ≥ 94%
Effective
On approval · forward decisions only

Decisions re-evaluated 4

  • Support agent / EnglishStayEligible
  • Support agent / SpanishStayEligible after quality review
  • FAQ fallback routeStayEligible
  • Release · GPT-5 → Claude Sonnet 5BlockedEligible
Historical replay
95.3% resolutionfloor 94%
Expected cost / resolution
−24%
Recommended next step
Add Meridian to the next approved canary

Version history

  1. v4Anthropic on AWS Bedrock added to permitted providersTomás Herrera · Security reviewjust now
  2. v3Spanish resolution floor raised 92% → 94%Meridian Health · Legal4w
  3. v2Refunds require human approvalMeridian Health · Security3mo

Your product changes.Your customer agreements change.Production should stay aligned with both.

Ask Synvolv

Ask the question you already have.

Your team should not need to know which dashboard contains the answer. Synvolv answers from the same customer-aware operating context and shows the evidence behind the answer.

Why did Acme’s support agent get expensive after v17?

Acme · support agent · since v17

Synvolv answered from your connected systems

Read 5 sources · Traces · GitHub · Evals · Billing · Customer records

Root cause found

Evidence 5

Refund-verification calls increased
1.2 → 3.5 / conversation
Traces
Change introduced in
PR #1816
GitHub
Resolution
+0.6 pt
Evals
Cost / resolution
$0.33 → $0.58
Billing
Also affected
4 customer accounts
Customer records

Suggested next step

Replay a one-retry limit against affected workflows.

Eligible first

“Roll it out to 5%.”5% of whom?

A percentage rollout tells you how much traffic to move. It does not tell you which customers should be inside that traffic. Synvolv resolves the eligible customer set first.

Eligible set

support-agent v17GPT-5 → Claude Sonnet 5evaluated 4 min ago

63 customers considered

27eligible

  • AcmeBusinessReplay: resolution +0.6 pt · quality floor met
  • NorthstarBusinessReplay: quality floor met · within cost limit
  • MonarchProQuality floor met · past included credits, cap applies
  • VelaStandardReplay: groundedness inside threshold · quality floor met

+ 23 more

8excluded by entitlement

  • HaldenEnterprisePlan entitles GPT-5 · v17 changes the model

+ 7 more

4excluded by policy

  • Meridian HealthEnterpriseBAA covers Azure OpenAI only

+ 3 more

2held for missing evidence

  • OrielStandardInsufficient replay coverage · not enough evidence it works for them

+ 1 more

22unaffected

No traffic on the changed workflow in the last 30 days

Release scope5% of traffic from 27 eligible customers · 36 customers unchanged

Support-agent traffic, one cell per account · coloured where v17 runs

Thenwith Synvolv

5% of eligible traffic.

Everyone in the 5% is eligible.

  • Meridian Healthnot in the 5%BAA covers Azure OpenAI only
  • Haldennot in the 5%plan entitles GPT-5
  • Orielnot in the 5%no evidence it works for them

Notwith a percentage rollout

5% of everybody.

14 accounts in the 5% are not eligible.

  • Meridian Healthin the 5%BAA covers Azure OpenAI only
  • Haldenin the 5%plan entitles GPT-5
  • Orielin the 5%no evidence it works for them

During the canary

One customer regressed. One customer rolled back.

  1. Every customer is watched on its own terms. During the canary each account is measured against the quality floor in its own record, because the average hides the one that is slipping.

  2. One customer slipped below its floor. Synvolv caught it in the first bad batch, because the floor was that customer's, not a global threshold.

  3. Only that customer went back. Its traffic returned to the previous version within minutes. The rest kept running, and never noticed.

Canary · support-agent v17

GPT-5 → Claude Sonnet 5 · 5% of traffic · running 41 min

Customers 5 of 41

Acmev170.94above floor0.88Healthy
Northstarv170.93above floor0.90Healthy
Monarchv170.95above floor0.85Healthy
Meridian——Not exposedExcluded by policy
Velav17rolled back to v160.81below floor0.90Groundedness regression detected

Activity

  1. 6 min agoSynvolv

    Continue canary26 customers stay on v17

  2. 7 min agoSynvolv

    Vela → previous versionRolled back to v16

  3. 8 min agoSynvolv

    Regression detected · Vela0.81 against a 0.90 floor

  4. 41 min agoPriya Raman

    Started canary5% of traffic · 27 eligible

Roll back the customer that regressed — not every customer that didn’t.

How far it goes

Start with answers.Add control after the answers earn your trust.

A new platform should not need your production traffic before it can prove that its decisions are useful.

Three stages. Each one gives Synvolv more reach and gives you a harder control in return: a decision you read, an action you approve, a limit enforced before the request runs. Each stage is your call. Stop at any of them.

  1. 01Shadow

    Reads

    You give
    Read access to your systems.
    You get
    The decision Synvolv would make, with the evidence behind it.

    Would release to 27 of 41 customers

    GPT-5 → Claude Sonnet 5 · 12,420 scenarios replayed

    Recommends only · nothing is applied

    Acme
    Move
    Meridian Health
    Stay · provider policy
    Vela
    Move · watch groundedness

    See what Synvolv would decide.

    Connect

    • GitHub
    • Production telemetry
    • Customer identity
    • Evaluation / outcome signals
    • Configuration history

    Then

    • Ask questions.
    • Review changes.
    • Replay historical scenarios.
    • See affected customers.
    • See recommended release scope.

    Read-only.No production write path required.

    after the answers earn your trust

  2. 02Approved actions

    Acts with approval

    You give
    Permission to act, one approval at a time.
    You get
    Canaries, pauses and rollbacks, through the tools you already run.

    Roll back Vela → v16?

    groundedness 0.81 · below Vela's floor of 0.90

    Nothing runs until someone approves

    Proposed by
    Synvolv · 2 min ago
    Scope
    Vela only · 26 customers continue
    Runs through
    your release controls

    Let the recommendation become an action.

    When your team is ready

    • Create customer-scoped canaries.
    • Pause a rollout.
    • Roll back an individual customer.
    • Apply approved policies.
    • Execute through supported gateways, release controls, and application hooks.

    Your existing infrastructure stays in place.

    only where exact control is worth it

  3. 03Synvolv Gateway

    In the request path

    You give
    The request path, for the workloads you choose.
    You get
    Hard limits enforced before a request runs.

    Daily agent budget $600

    used $572 · forecast today $941

    Enforced in the request path

    Approved response
    Fallback after $600
    Applies to
    Acme · support agent
    Checked
    before every request

    Put Synvolv in the request path only where exact control is worth it.

    For workloads that need hard runtime enforcement

    • Per-customer routing
    • Hard tenant budgets
    • Pre-request spend reservation
    • Entitlement enforcement
    • Provider / model restrictions
    • Tool-call limits
    • Agent-step limits
    • Loop protection
    • Immediate fallbacks
    • Usage + cost reconciliation

Test your next AI change

Runtime control

“Why did this tenant burn $900 before lunch?”

A successful release is not the end of the job.

One tenant can become disproportionately expensive without the global dashboard looking alarming.

Acme · the morning, five minutes at a timethis morning

  1. Agents loop.
  2. Usage spikes.
  3. A fallback starts using the expensive model.
  4. A tool gets called five times when one was enough.
  5. A customer burns through included usage.
  1. Agents loop.
  2. Usage spikes.
  3. A fallback starts using the expensive model.
  4. A tool gets called five times when one was enough.
  5. A customer burns through included usage.
  6. Gateway · $600

Fallback after $600Cap expensive workflowRequire approvalContinue + generate overage event

Used $572Budget $600Forecast today $941

With Synvolv Gateway, those customer-level limits can be enforced before execution — not discovered after the invoice arrives.

Customer cost-to-serve

Know what every customer costs you to serve — and what made them expensive.

Your provider invoice tells you what you spent.A token chart tells you usage grew.Neither tells you why one customer stopped making economic sense.

Synvolv connects cost back to the customer · the workflow · the agent version · the tool behavior · the release · the outcome.

When commercial context is connected, add revenue · included usage · credits · contract terms · contribution margin.

Acme

Revenue
$4,000 / month
Agent cost
$1,940 / month
Contribution marginMargin
41% / 55% target
Target
55%

What changed? Support Agent v17 · Refund-tool calls / resolution 1.2 → 3.5 · Cost / successful resolution $0.33 → $0.58

v16 before the release

  • Retrieve account$0.04
  • Classify intent$0.05
  • Refund tool · call 11.2 calls on average$0.10
  • Confirm outcome$0.08
  • Close$0.06

Cost / successful resolution$0.33

v17 now

On v16 this resolution cost $0.33.

  • Retrieve account$0.04
  • Classify intent$0.05
  • Refund tool · call 1$0.10
  • Refund tool · call 2$0.10
  • Refund tool · call 3$0.10
  • Refund tool · call 4half of resolutions$0.05
  • Confirm outcome$0.08
  • Close$0.06

Cost / successful resolution$0.58

What can we do?

Choose an option to see what it does to this resolution.

Don’t just find the expensive customer.Find the behavior making them expensive.

One context

The systems you already use know pieces of the answer.

  • Your evaluator can measure behavior.
  • Your gateway can execute a route.
  • Your release system can expose a variation.
  • Your billing platform can meter usage.
  • Your CRM knows the customer.
  • Your contract system stores the terms.
  • Your finance tooling sees the dollars.

Synvolv resolves those facts into the customer-level operating decision.

  • Should this customer receive this behavior?
  • What changes for them?
  • Is it permitted?
  • Are they entitled to it?
  • What will it cost to serve them?
  • What should happen next?

And after the action:Did reality match the plan?

The operating loop

One customer context across the operating loop.

  1. Change

    What changed in the model, prompt, retrieval, tool, workflow, code, or customer rule?

  2. Understand

    Which customers and workloads does it touch?

  3. Test

    What changes for those customers?

  4. Resolve

    Who should receive it now?

  5. Plan

    What exactly should happen?

  6. Approve

    Who must sign off?

  7. Act

    Release, route, limit, pause, or enforce.

  8. Verify

    What actually happened?

  9. Learn

    Use the result the next time.

Review, Ask, Release, Runtime, Gateway, and Economics are not six unrelated products. They are different moments where the same customer operating context gets used.

Find out who it changes before your customers do.

Start with the change your team is already debating.

Bring any of theseA model migration.A prompt edit.A retrieval change.A new tool.A workflow update.A routing change.A customer-policy change.An AI-related pull request.

You do not have to trust Synvolv with production to see whether the answer is useful. Start in shadow. Add control when you are ready.