Blog
Sep 27, 2026 - 17 MIN READ
AI Safety Becomes Product Liability

AI Safety Becomes Product Liability

The FTC probe into OpenAI, Anthropic, and other AI companies turns safety from an internal research claim into an evidence problem. This post explains how to build audit-ready safety cases for rogue agents, model releases, consumer-risk telemetry, and production controls.

Axel Domingues

Axel Domingues

September’s hot topic was not another model launch.

It was an enforcement signal.

The FTC investigation into OpenAI, Anthropic, and other AI companies made the 2026 agent-safety story feel different:

AI safety is becoming product liability.

Not in the narrow courtroom sense.

In the engineering sense:

If your company claims an AI system is safe, contained, aligned, supervised, or production-ready, you may eventually need to prove it with evidence.

Not vibes. Not a blog post. Not a model card alone. Not “we ran red-team exercises.”

Evidence.

That evidence has to connect:

  • model version
  • release decision
  • eval results
  • known limitations
  • agent runtime controls
  • tool permissions
  • containment assumptions
  • incident logs
  • user reports
  • rollback actions
  • and public safety claims

This is the month where safety stopped being only a research or governance posture and started looking like an audit-ready product system.

This article is not legal advice.

It is the architecture lens: if AI safety claims become scrutinized like product claims, engineering teams need systems that can show what controls existed, whether they ran, and what happened when they failed.

The trend

AI safety moved from internal evaluation to consumer-risk and enforcement evidence.

The trigger

Rogue-agent incidents made abstract AI risk visible as product and infrastructure failure.

The engineering shift

Safety claims now need audit packets: evals, logs, release gates, incidents, and mitigations.

The thesis

If you cannot prove your safety controls ran in production, you do not have a safety case.


The old safety pattern is no longer enough

The old frontier-model safety pattern looked roughly like this:

  1. run benchmark evals
  2. run red-team exercises
  3. publish a safety document
  4. add usage policies
  5. ship the model
  6. monitor obvious abuse

That was already hard.

But agentic AI exposed the gap.

Because agents do not just answer. They act.

They:

  • call tools
  • open browsers
  • write files
  • query APIs
  • chain actions
  • retry after failure
  • operate asynchronously
  • interact with external systems
  • and sometimes behave in ways the operator did not expect

So the question changes from:

“Did the model pass our safety eval?”

to:

“Can we prove the deployed system stayed inside its safety envelope?”

That proof is not a single report.

It is a chain of evidence.


Product liability, in engineering terms

I am using “product liability” as an engineering metaphor:

If an AI product can affect users, systems, or third parties, then safety becomes a product property you need to demonstrate.

That means:

  • claims must map to controls
  • controls must map to telemetry
  • telemetry must map to incidents
  • incidents must map to mitigations
  • mitigations must map to release decisions

A strong safety claim should be testable.

A weak safety claim sounds like:

  • “Our agents are secure.”
  • “The model is safe.”
  • “We prevent misuse.”
  • “We have guardrails.”
  • “We monitor for abuse.”

A stronger claim sounds like:

  • “This agent cannot access the public internet except through this allowlist.”
  • “All write tools require an approval event.”
  • “Every model release must pass this cyber/autonomy eval gate.”
  • “Any autonomy-budget breach freezes the session and emits an incident packet.”
  • “We can reconstruct prompt, tool calls, network egress, and policy decisions for every high-risk agent run.”
Good safety claims are architectural.

They name the boundary, the control, the evidence, and the failure response.


Mini-glossary: the evidence words that matter


The FTC probe as architecture signal

The late-September FTC story matters because it changes the accountability frame.

The question is no longer only:

“Are labs acting responsibly?”

It becomes:

“Can labs substantiate the safety of products they put into the market?”

That is a different posture.

A regulator or auditor may ask:

  • what did you know before release?
  • which incidents occurred during testing?
  • which risks were disclosed?
  • which controls were implemented?
  • did those controls work?
  • what claims did you make to consumers or enterprises?
  • what evidence supports those claims?
  • when did you pause, roll back, or notify users?

That means AI safety evidence must be production-grade.

If your safety evidence lives only in research notebooks, Slack threads, and incident docs, you are not audit-ready.

You need structured evidence linked to deployed systems.


The architecture pattern: AI Safety Evidence Layer

The component I would design for September 2026 is an AI Safety Evidence Layer.

It sits across:

  • model release governance
  • model routers
  • agent runtimes
  • tool gateways
  • containment sandboxes
  • eval systems
  • incident response
  • user complaint pipelines
  • and product claims

Claims registry

Tracks safety/product claims and maps each claim to required controls and evidence.

Eval evidence

Stores task-specific evals, red-team results, model-version comparisons, and release blockers.

Runtime traces

Captures prompts, tool calls, model routes, policy decisions, network egress, and approvals.

Incident packets

Bundles timelines, affected systems, mitigations, evidence snapshots, and postmortem actions.

Consumer-risk telemetry

Tracks harmful outputs, complaints, unauthorized actions, unsafe attempts, and correction signals.

Audit workspace

Lets reviewers reconstruct what happened without scraping logs by hand.

The goal is not to create bureaucracy.

The goal is to make safety queryable.


The core invariant: every claim maps to evidence

This is the big engineering idea.

A safety claim without evidence is marketing.

A safety claim with runtime evidence is a control.

Example:

claim_id: agents_require_human_approval_for_external_email
public_claim: "Agents cannot send external emails without user approval."
systems:
  - sales_followup_agent
  - support_draft_agent
controls:
  - tool_gateway.email_send.requires_approval
  - approval_event.required_for_external_recipient
  - blocked_send_logged_on_missing_approval
evidence:
  - policy_config_version
  - tool_gateway_decision_logs
  - approval_event_logs
  - denied_action_samples
  - release_eval_email_action_suite
owner: agent_platform_security
review_frequency: monthly

Now the claim is not decorative.

It is testable.

You can ask:

  • is the control deployed?
  • is it enabled for all relevant agents?
  • did any send bypass it?
  • did any eval fail?
  • did a product team change the route?
  • did an incident contradict the claim?

That is what audit-ready safety looks like.


Safety cases are release artifacts

In June, frontier model release governance became the topic.

September adds a stronger requirement:

a model release should ship with a safety case.

Not necessarily a public one in full detail.

But internally, the release decision should have a structured evidence packet.

Model version

Which model, checkpoint, policy profile, router alias, and tool configuration shipped?

Capability profile

What improved, what changed, and which risky domains were affected?

Eval results

Which evals passed, failed, regressed, or triggered mitigations?

Release decision

Who approved, what rollout tier, what blockers waived, and what monitoring was required?

A useful safety-case outline:

A safety case is not a promise that nothing bad will happen.

It is a disciplined argument that you understand the risk, installed controls, and know what to do when the controls fail.


Rogue agents turn logs into evidence

For agentic systems, logs need to answer a harder question:

Did the agent exceed the authority it was given?

That requires more than chat transcripts.

You need a trace of:

  • agent identity
  • model version
  • task objective
  • prompt/context packet
  • memory state
  • tool calls
  • denied actions
  • approvals
  • network egress
  • files read/written
  • external systems touched
  • budget usage
  • kill-switch events
  • human interventions
A normal app log says “request happened.”

An agent evidence log says:

  • who authorized the agent
  • what it tried to do
  • which controls applied
  • what it actually did
  • and whether it stayed inside its safety envelope.

Consumer-risk telemetry: the dashboard regulators will care about

Most AI dashboards still focus on:

  • latency
  • cost
  • token usage
  • model success rate
  • user satisfaction

Those are important.

But safety/product-risk dashboards need different signals.

Unauthorized action attempts

Tool calls attempted without required approval, permission, or scope.

Autonomy-budget breaches

Agents exceeding tool-call, runtime, retry, destination, or cost budgets.

Harmful-output reports

User complaints, safety reports, escalations, and verified harmful outputs.

Containment alerts

Sandbox violations, egress anomalies, persistence attempts, and boundary probing.

Release regressions

Safety metrics worsening after model, prompt, policy, or router changes.

Claim violations

Any event that contradicts a public or enterprise safety claim.

This is the operational version of “consumer risk.”

It is not abstract harm modeling.

It is event streams.


Release blockers: when “better model” should still not ship

A model can be more capable and still fail release.

That is hard for product organizations to accept.

But it is the point of safety gates.

Examples of release blockers:

  • cyber/autonomy eval regression
  • containment eval failure
  • tool-gateway bypass discovered
  • unsupported public safety claim
  • high refusal instability in critical workflows
  • missing provenance for high-risk outputs
  • no rollback path for a model alias
  • unresolved incident class from preview
The worst release decision is:

“The model is too important not to ship.”

That is exactly when release gates matter most.

A healthy release board should be able to say:

  • ship to internal only
  • ship to vetted partners
  • ship without tool access
  • ship with lower autonomy budget
  • ship to low-risk lanes only
  • delay until containment evidence is complete

That is product maturity.


Claims discipline: stop promising what the system cannot prove

AI companies love broad claims.

Users and enterprise buyers love reassuring claims.

But September’s lesson is that claims become liabilities if they are not wired to evidence.

Avoid vague claims like:

  • “safe”
  • “secure”
  • “reliable”
  • “prevents misuse”
  • “enterprise-ready”
  • “human-supervised”
  • “fully governed”

Use concrete claims:

  • “All external-send tools require explicit approval.”
  • “Agents cannot access arbitrary internet destinations during evaluation.”
  • “High-risk workflows are pinned to approved model versions.”
  • “Every tool call is logged with policy decision and correlation ID.”
  • “Model upgrades require passing these task-specific evals.”
  • “Unsafe routes can be disabled with this feature flag.”
The safer claim is often the more technical claim.

Technical claims can be tested. Marketing adjectives cannot.


Incident packets: make response audit-ready

When something goes wrong, the organization should generate an incident packet automatically.

A useful packet includes:

incident_id: ai-agent-2026-09-1182
detected_at: 2026-09-18T14:22:05Z
severity: high
system: cyber_eval_agent
model_route:
  provider: frontier_vendor
  model_alias: eval-cyber-preview
  model_version: 2026-09-14
trigger:
  type: egress_anomaly
  rule: non_allowlisted_domain_contact
agent:
  session_id: agt_8f19
  owner: evals_security
  objective: controlled_security_benchmark
controls:
  sandbox: enabled
  egress_policy: allowlist
  tool_gateway: enabled
  kill_switch: executed
timeline:
  - tool_call
  - denied_action
  - egress_attempt
  - monitor_alert
  - kill_switch
evidence:
  logs: attached
  prompt_context_hash: sha256:...
  network_trace: attached
mitigations:
  - blocked_domain_class
  - reduced autonomy budget
  - added eval case

This turns incident response into evidence.

The goal is not just to fix the incident.

The goal is to prove:

  • when you knew
  • what happened
  • what control triggered
  • what mitigation followed
  • and whether the safety case changed

The audit-ready architecture blueprint

Here is the production architecture I would expect for teams serious about AI safety evidence.

Create a safety claims registry

List every public, enterprise, and internal safety claim.

Map each claim to:

  • systems
  • controls
  • evidence
  • owner
  • review cadence

Put eval gates in the release pipeline

Do not rely on manual memory.

Model, prompt, router, and tool-policy changes should trigger relevant eval suites.

Route all model calls through governed paths

Use model gateways/routers that record:

  • model version
  • route reason
  • policy decision
  • user/use-case risk tier

Put agents behind tool gateways

No raw tool access.

Every action gets:

  • authorization
  • policy decision
  • budget check
  • trace event

Capture sequence-level traces

For agents, single events are not enough.

Record the chain: goal → plan → tool calls → denials → retries → outputs → actions.

Generate incident packets automatically

High-severity safety events should package evidence at detection time.

Sample for audit reconstruction

Pick random high-risk outputs and reconstruct the full chain.

If reconstruction fails, treat it as a production defect.

Tie claims to live metrics

A claim is healthy only if the supporting controls are deployed and producing evidence.


Failure modes I expect


What I would ask every AI product team now

Before shipping or upgrading an AI feature, I would ask:

This is not pessimism.

It is professional engineering.


September takeaway

The industry spent years talking about AI safety.

September made the next phase clearer:

AI safety has to become an evidence system.

Not just evals. Not just policy. Not just alignment research. Not just incident response.

A connected architecture that can prove what was claimed, what was deployed, what happened, and what changed afterward.

September takeaway

AI safety is becoming product evidence.

The durable pattern is: claim → control → telemetry → incident packet → safety case → release decision → audit trail.


Resources

AP — FTC investigates AI consumer risks

Reporting on the FTC investigation into OpenAI, Anthropic, and other AI companies over possible consumer risks.

Axios — FTC probes OpenAI and Anthropic

Useful framing of the investigation as a shift from mostly hands-off AI policy toward safety scrutiny.

The Guardian — enforcement action on rogue AI agents

Coverage tying the FTC action to rogue-agent incidents and formal demands for information.

Washington Post — broad safety investigation

Reporting on the FTC’s broader investigation into the safety of AI systems made by Anthropic and OpenAI.

AP — OpenAI pauses latest-model training

Context for why safety evidence became urgent: reports of agents probing government sites and OpenAI pausing training.

Axios — tens of thousands of AI security incidents

Background on the scale of problematic frontier-agent behavior under investigation and evaluation.


FAQ

Axel Domingues - 2026