[{"data":1,"prerenderedAt":2711},["ShallowReactive",2],{"navigation":3,"/blog/ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases":526,"/blog/ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases-surround":2707},[4],{"title":5,"path":6,"stem":7,"children":8,"page":525},"Blog","/blog","blog",[9,13,17,21,25,29,33,37,41,45,49,53,57,61,65,69,73,77,81,85,89,93,97,101,105,109,113,117,121,125,129,133,137,141,145,149,153,157,161,165,169,173,177,181,185,189,193,197,201,205,209,213,217,221,225,229,233,237,241,245,249,253,257,261,265,269,273,277,281,285,289,293,297,301,305,309,313,317,321,325,329,333,337,341,345,349,353,357,361,365,369,373,377,381,385,389,393,397,401,405,409,413,417,421,425,429,433,437,441,445,449,453,457,461,465,469,473,477,481,485,489,493,497,501,505,509,513,517,521],{"title":10,"path":11,"stem":12},"Activation Functions Are Not a Detail - ReLU Changed Everything","/blog/activation-functions-are-not-a-detail","blog/activation-functions-are-not-a-detail",{"title":14,"path":15,"stem":16},"Actor-Critic - The First Time RL Feels Trainable","/blog/actor-critic-the-first-time-rl-feels-trainable","blog/actor-critic-the-first-time-rl-feels-trainable",{"title":18,"path":19,"stem":20},"Agent Evals as CI - From Prompt Tests to Scenario Harnesses and Red Teams","/blog/agent-evals-as-ci-from-prompt-tests-to-scenario-harnesses-and-red-teams","blog/agent-evals-as-ci-from-prompt-tests-to-scenario-harnesses-and-red-teams",{"title":22,"path":23,"stem":24},"Agent Runtimes Emerge: SDKs, orchestration primitives, and observability","/blog/agent-runtimes-emerge-sdks-orchestration-primitives-and-observability","blog/agent-runtimes-emerge-sdks-orchestration-primitives-and-observability",{"title":26,"path":27,"stem":28},"Agentic AI Is Becoming a Cybersecurity Problem","/blog/agentic-ai-is-becoming-a-cybersecurity-problem","blog/agentic-ai-is-becoming-a-cybersecurity-problem",{"title":30,"path":31,"stem":32},"Agents as Distributed Systems: outbox, sagas, and “eventually correct” workflows","/blog/agents-as-distributed-systems-outbox-sagas-eventually-correct-workflows","blog/agents-as-distributed-systems-outbox-sagas-eventually-correct-workflows",{"title":34,"path":35,"stem":36},"AI Compliance Becomes Runtime Architecture","/blog/ai-compliance-becomes-runtime-architecture-gpai-transparency-auditability","blog/ai-compliance-becomes-runtime-architecture-gpai-transparency-auditability",{"title":38,"path":39,"stem":40},"AI Safety Becomes Product Liability","/blog/ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases","blog/ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases",{"title":42,"path":43,"stem":44},"AJAX → Fetch → GraphQL → tRPC: Choosing Your Data Boundary","/blog/ajax-fetch-graphql-trpc-choosing-your-data-boundary","blog/ajax-fetch-graphql-trpc-choosing-your-data-boundary",{"title":46,"path":47,"stem":48},"Exercise 8 + Course Wrap - Anomaly Detection & Recommenders (and My Next Steps)","/blog/anomaly-detection-and-recommenders","blog/anomaly-detection-and-recommenders",{"title":50,"path":51,"stem":52},"API Evolution at Scale: Compatibility, Contracts, and Consumer-Driven Testing","/blog/api-evolution-at-scale-compatibility-contracts-consumer-driven-testing","blog/api-evolution-at-scale-compatibility-contracts-consumer-driven-testing",{"title":54,"path":55,"stem":56},"Backends: Frameworks Don’t Matter Until They Do (Node, Java, .NET, Go, Python)","/blog/backends-frameworks-dont-matter-until-they-do","blog/backends-frameworks-dont-matter-until-they-do",{"title":58,"path":59,"stem":60},"Backpropagation Demystified - It’s Just the Chain Rule (But Applied Ruthlessly)","/blog/backpropagation-demystified","blog/backpropagation-demystified",{"title":62,"path":63,"stem":64},"Bandits - The First Honest RL Problem","/blog/bandits-the-first-honest-rl-problem","blog/bandits-the-first-honest-rl-problem",{"title":66,"path":67,"stem":68},"Batch Training & Evaluation Again: Promising Results That Survive Scrutiny","/blog/batch-training-evaluation-again-promising-results-that-survive-scrutiny","blog/batch-training-evaluation-again-promising-results-that-survive-scrutiny",{"title":70,"path":71,"stem":72},"bitmex-gym - The Baseline Trading Environment (Where Cheating Starts)","/blog/bitmex-gym-baseline-trading-environment-where-cheating-starts","blog/bitmex-gym-baseline-trading-environment-where-cheating-starts",{"title":74,"path":75,"stem":76},"bitmex-management-gym: Position Sizing and the First Risk-Aware Agent","/blog/bitmex-management-gym-position-sizing-first-risk-aware-agent","blog/bitmex-management-gym-position-sizing-first-risk-aware-agent",{"title":78,"path":79,"stem":80},"Browser Reality: The Event Loop, Rendering, and Why UX Bugs Look Like Backend Bugs","/blog/browser-reality-event-loop-rendering-ux-bugs-backend-bugs","blog/browser-reality-event-loop-rendering-ux-bugs-backend-bugs",{"title":82,"path":83,"stem":84},"Caching Without Folklore: Redis, CDNs, and the Two Hard Things","/blog/caching-without-folklore-redis-cdns-and-the-two-hard-things","blog/caching-without-folklore-redis-cdns-and-the-two-hard-things",{"title":86,"path":87,"stem":88},"Capstone: Build a System That Can Survive (Reference Architecture + Decision Log)","/blog/capstone-build-a-system-that-can-survive","blog/capstone-build-a-system-that-can-survive",{"title":90,"path":91,"stem":92},"Chappie Wiring From Trained Policy to Running Process","/blog/chappie-wiring-from-trained-policy-to-running-process","blog/chappie-wiring-from-trained-policy-to-running-process",{"title":94,"path":95,"stem":96},"CI/CD as Architecture: Testing Pyramids, Pipelines, and Rollout Safety","/blog/ci-cd-as-architecture-testing-pyramids-pipelines-rollout-safety","blog/ci-cd-as-architecture-testing-pyramids-pipelines-rollout-safety",{"title":98,"path":99,"stem":100},"Cloud Infrastructure Without the Fanaticism: IaaS, PaaS, Serverless, Kubernetes","/blog/cloud-infrastructure-without-the-religion","blog/cloud-infrastructure-without-the-religion",{"title":102,"path":103,"stem":104},"Computer-Use Agents in Production: sandboxes, VMs, and UI-action safety","/blog/computer-use-agents-in-production-sandboxes-vms-ui-action-safety","blog/computer-use-agents-in-production-sandboxes-vms-ui-action-safety",{"title":106,"path":107,"stem":108},"Constraints That Teach: Risk Caps, Timeouts, and Surviving Bad Regimes","/blog/constraints-that-teach-risk-caps-timeouts-surviving-bad-regimes","blog/constraints-that-teach-risk-caps-timeouts-surviving-bad-regimes",{"title":110,"path":111,"stem":112},"Containers, Docker, and the Discipline of Reproducibility","/blog/containers-docker-and-the-discipline-of-reproducibility","blog/containers-docker-and-the-discipline-of-reproducibility",{"title":114,"path":115,"stem":116},"Context Assembly as a Subsystem: Summaries, State, and Token Budgets","/blog/context-assembly-as-a-subsystem-summaries-state-and-token-budgets","blog/context-assembly-as-a-subsystem-summaries-state-and-token-budgets",{"title":118,"path":119,"stem":120},"Continuous Control - DDPG and the Seduction of Off-Policy","/blog/continuous-control-ddpg-and-the-seduction-of-off-policy","blog/continuous-control-ddpg-and-the-seduction-of-off-policy",{"title":122,"path":123,"stem":124},"Convolutions - Why CNNs See the World Differently","/blog/convolutions-why-cnns-see-the-world-differently","blog/convolutions-why-cnns-see-the-world-differently",{"title":126,"path":127,"stem":128},"Cost as a First-Class Constraint: FinOps for Architects","/blog/cost-as-a-first-class-constraint-finops-for-architects","blog/cost-as-a-first-class-constraint-finops-for-architects",{"title":130,"path":131,"stem":132},"DALL·E: How Text Became Images (and Why It Changed Everything)","/blog/dalle-how-text-became-images-and-why-it-changed-everything","blog/dalle-how-text-became-images-and-why-it-changed-everything",{"title":134,"path":135,"stem":136},"Data Engineering for Product Teams: OLTP vs OLAP, Streaming, and Truth","/blog/data-engineering-for-product-teams-oltp-vs-olap-streaming-and-truth","blog/data-engineering-for-product-teams-oltp-vs-olap-streaming-and-truth",{"title":138,"path":139,"stem":140},"Data Stores 101 for Architects: SQL, NoSQL, and the Shape of Consistency","/blog/data-stores-101-for-architects-sql-nosql-consistency","blog/data-stores-101-for-architects-sql-nosql-consistency",{"title":142,"path":143,"stem":144},"Dataset Reality — HDF5 Schema, Missing Data, and “Don’t Lie to Yourself” Rules","/blog/dataset-reality-hdf5-schema-missing-data","blog/dataset-reality-hdf5-schema-missing-data",{"title":146,"path":147,"stem":148},"Exercise 5 - Debugging ML (Bias/Variance, Learning Curves, and What to Try Next)","/blog/debugging-ml-bias-variance","blog/debugging-ml-bias-variance",{"title":150,"path":151,"stem":152},"Deep Q-Learning - My First Real Baselines Month","/blog/deep-q-learning-my-first-real-baselines-month","blog/deep-q-learning-my-first-real-baselines-month",{"title":154,"path":155,"stem":156},"Deep Silos in RL: Architecture as Stability (and the First LSTM Variant)","/blog/deep-silos-in-rl-architecture-as-stability","blog/deep-silos-in-rl-architecture-as-stability",{"title":158,"path":159,"stem":160},"Deep Silos - Representation Learning That Respects Feature Families","/blog/deep-silos-representation-learning-feature-families","blog/deep-silos-representation-learning-feature-families",{"title":162,"path":163,"stem":164},"Defining Alpha Without Cheating - Look-Ahead Labels and Leakage Traps","/blog/defining-alpha-without-cheating","blog/defining-alpha-without-cheating",{"title":166,"path":167,"stem":168},"Dissecting ChatGPT: The Product Architecture Around the Model","/blog/dissecting-chatgpt-the-product-architecture-around-the-model","blog/dissecting-chatgpt-the-product-architecture-around-the-model",{"title":170,"path":171,"stem":172},"Distributed Data: Transactions, Outbox, Sagas, and “Eventually Correct”","/blog/distributed-data-transactions-outbox-sagas-eventually-correct","blog/distributed-data-transactions-outbox-sagas-eventually-correct",{"title":174,"path":175,"stem":176},"Why Sequences Break Everything - Enter Recurrent Neural Networks","/blog/enter-recurrent-neural-networks","blog/enter-recurrent-neural-networks",{"title":178,"path":179,"stem":180},"Evaluation Discipline - Walk-Forward Backtesting Inside the Gym","/blog/evaluation-discipline-walk-forward-backtesting-inside-gym","blog/evaluation-discipline-walk-forward-backtesting-inside-gym",{"title":182,"path":183,"stem":184},"Feature Engineering, But Make It Microstructure: Liquidity Created/Removed","/blog/feature-engineering-microstructure-liquidity-created-removed","blog/feature-engineering-microstructure-liquidity-created-removed",{"title":186,"path":187,"stem":188},"First Live Runs - Small Size, Big Lessons","/blog/first-live-runs-small-size-big-lessons","blog/first-live-runs-small-size-big-lessons",{"title":190,"path":191,"stem":192},"From Logistic Regression to Neurons - Rebuilding Intuition from the Perceptron","/blog/from-logistic-regression-to-neurons","blog/from-logistic-regression-to-neurons",{"title":194,"path":195,"stem":196},"From Microstructure to Features - What the Model Will See","/blog/from-microstructure-to-features-what-the-model-will-see","blog/from-microstructure-to-features-what-the-model-will-see",{"title":198,"path":199,"stem":200},"From Classical ML to Deep Learning - What Actually Changed (and What Didn’t) (and My Next Steps)","/blog/from-ml-to-deep-learning-retrospective","blog/from-ml-to-deep-learning-retrospective",{"title":202,"path":203,"stem":204},"From Prediction to Decision - Designing the Trading Environment Contract","/blog/from-prediction-to-decision-trading-environment-contract","blog/from-prediction-to-decision-trading-environment-contract",{"title":206,"path":207,"stem":208},"From Research Rig to System: 2020 Postmortem and the Real Amazing Result","/blog/from-research-rig-to-system-2020-postmortem","blog/from-research-rig-to-system-2020-postmortem",{"title":210,"path":211,"stem":212},"Frontend Systems: Routing, State, Forms, and the “Boring Stack” That Scales","/blog/frontend-systems-routing-state-forms-boring-stack","blog/frontend-systems-routing-state-forms-boring-stack",{"title":214,"path":215,"stem":216},"Frontier Model Release Governance","/blog/frontier-model-release-governance-national-security-workflow","blog/frontier-model-release-governance-national-security-workflow",{"title":218,"path":219,"stem":220},"Function Approximation - The Day RL Stopped Being Stable","/blog/function-approximation-the-day-rl-stopped-being-stable","blog/function-approximation-the-day-rl-stopped-being-stable",{"title":222,"path":223,"stem":224},"GPAI Obligations Begin: What Changes for Model Providers and Enterprises","/blog/gpai-obligations-begin-what-changes-for-model-providers-and-enterprises","blog/gpai-obligations-begin-what-changes-for-model-providers-and-enterprises",{"title":226,"path":227,"stem":228},"Hallucinations: A Probabilistic Failure Mode, Not a Moral Defect","/blog/hallucinations-a-probabilistic-failure-mode-not-a-moral-defect","blog/hallucinations-a-probabilistic-failure-mode-not-a-moral-defect",{"title":230,"path":231,"stem":232},"HTTP as a Distributed Systems API (Without the Buzzwords)","/blog/http-as-a-distributed-systems-api-without-the-buzzwords","blog/http-as-a-distributed-systems-api-without-the-buzzwords",{"title":234,"path":235,"stem":236},"Imitation Learning - GAIL and the Strange Feeling of Learning From Experts","/blog/imitation-learning-gail-and-learning-from-experts","blog/imitation-learning-gail-and-learning-from-experts",{"title":238,"path":239,"stem":240},"Incident Response and Resilience: Designing for Failure, Not Hope","/blog/incident-response-and-resilience-designing-for-failure-not-hope","blog/incident-response-and-resilience-designing-for-failure-not-hope",{"title":242,"path":243,"stem":244},"Initialization, Scale, and the Fragility of Deep Networks","/blog/initialization-scale-fragility-of-deep-networks","blog/initialization-scale-fragility-of-deep-networks",{"title":246,"path":247,"stem":248},"Instruction Tuning: Turning a Completion Engine into an Assistant","/blog/instruction-tuning-turning-a-completion-engine-into-an-assistant","blog/instruction-tuning-turning-a-completion-engine-into-an-assistant",{"title":250,"path":251,"stem":252},"Exercise 3 - One-vs-All + Intro to Neural Networks (Handwritten Digits!)","/blog/intro-to-neural-networks","blog/intro-to-neural-networks",{"title":254,"path":255,"stem":256},"Exercise 1 - Linear Regression From Scratch","/blog/linear-regression-from-scratch","blog/linear-regression-from-scratch",{"title":258,"path":259,"stem":260},"Linear Regression With Multiple Variables (and Why Vectorization Matters)","/blog/linear-regression-with-multiple-vars","blog/linear-regression-with-multiple-vars",{"title":262,"path":263,"stem":264},"Live Alpha Monitoring - When the Market Talks Back","/blog/live-alpha-monitoring-when-market-talks-back","blog/live-alpha-monitoring-when-market-talks-back",{"title":266,"path":267,"stem":268},"Exercise 2 - Logistic Regression for Classification (My First Real Classifier)","/blog/logistic-regression-for-classification","blog/logistic-regression-for-classification",{"title":270,"path":271,"stem":272},"Long Context Isn’t Memory: When to Stuff, When to Retrieve","/blog/long-context-isnt-memory-when-to-stuff-when-to-retrieve","blog/long-context-isnt-memory-when-to-stuff-when-to-retrieve",{"title":274,"path":275,"stem":276},"LSTMs - Engineering Memory into the Network","/blog/lstms-engineering-memory-into-the-network","blog/lstms-engineering-memory-into-the-network",{"title":278,"path":279,"stem":280},"Maker Trades as a Strategy: When Fees Become a Reward Signal","/blog/maker-trades-fees-reward-signal","blog/maker-trades-fees-reward-signal",{"title":282,"path":283,"stem":284},"Microservices vs Modular Monolith: The “When” and the “How”","/blog/microservices-vs-modular-monolith-when-and-how","blog/microservices-vs-modular-monolith-when-and-how",{"title":286,"path":287,"stem":288},"Midjourney and the Product Loop: Why Some Generators Feel Magical","/blog/midjourney-and-the-product-loop-why-some-generators-feel-magical","blog/midjourney-and-the-product-loop-why-some-generators-feel-magical",{"title":290,"path":291,"stem":292},"Model Selection Becomes Architecture: Routing, Budgets, and Capability Tiers","/blog/model-selection-becomes-architecture-routing-budgets-and-capability-tiers","blog/model-selection-becomes-architecture-routing-budgets-and-capability-tiers",{"title":294,"path":295,"stem":296},"Multi-Agent Systems Without Chaos: supervisors, specialists, and coordination contracts","/blog/multi-agent-systems-without-chaos-supervisors-specialists-and-coordination-contracts","blog/multi-agent-systems-without-chaos-supervisors-specialists-and-coordination-contracts",{"title":298,"path":299,"stem":300},"Multimodal Changes UX: designing text+vision+audio systems","/blog/multimodal-changes-ux-designing-text-vision-audio-systems","blog/multimodal-changes-ux-designing-text-vision-audio-systems",{"title":302,"path":303,"stem":304},"Exercise 4 - Neural Networks Learning (Backpropagation Without Tears)","/blog/neural-networks-learning-backpropagation","blog/neural-networks-learning-backpropagation",{"title":306,"path":307,"stem":308},"Normal Equation vs Gradient Descent (Choosing Tools Like an Engineer)","/blog/normal-equation-vs-gradient-descent","blog/normal-equation-vs-gradient-descent",{"title":310,"path":311,"stem":312},"Normalization Is a Deployment Problem - Mean/Sigma and Index Diff","/blog/normalization-is-a-deployment-problem","blog/normalization-is-a-deployment-problem",{"title":314,"path":315,"stem":316},"Observability that Works: Logs, Metrics, Traces, and SLO Thinking","/blog/observability-that-works-logs-metrics-traces-and-slo-thinking","blog/observability-that-works-logs-metrics-traces-and-slo-thinking",{"title":318,"path":319,"stem":320},"Open Weights in Production: evaluation, licensing, and guardrails","/blog/open-weights-in-production-evaluation-licensing-and-guardrails","blog/open-weights-in-production-evaluation-licensing-and-guardrails",{"title":322,"path":323,"stem":324},"OpenClaw: A Viral Agent, a Skills Ecosystem, and the Supply-Chain Reality Check","/blog/openclaw-a-viral-agent-and-the-supply-chain-reality-check","blog/openclaw-a-viral-agent-and-the-supply-chain-reality-check",{"title":326,"path":327,"stem":328},"Optimization Got Real - Momentum, Learning Rates, and Why Plain Gradient Descent Wasn’t Enough","/blog/optimization-got-real-momentumand-learning-rates","blog/optimization-got-real-momentumand-learning-rates",{"title":330,"path":331,"stem":332},"Order Books Are the Battlefield - Matching Engines in Plain English","/blog/order-books-are-the-battlefield","blog/order-books-are-the-battlefield",{"title":334,"path":335,"stem":336},"Performance Engineering End-to-End: From TTFB to Tail Latency","/blog/performance-engineering-end-to-end-from-ttfb-to-tail-latency","blog/performance-engineering-end-to-end-from-ttfb-to-tail-latency",{"title":338,"path":339,"stem":340},"Policy Gradients - Learning Without a Value Crutch","/blog/policy-gradients-learning-without-a-value-crutch","blog/policy-gradients-learning-without-a-value-crutch",{"title":342,"path":343,"stem":344},"Pooling, Hierarchies, and What CNNs Are Really Learning","/blog/pooling-hierarchies-and-cnns","blog/pooling-hierarchies-and-cnns",{"title":346,"path":347,"stem":348},"Pretraining Is Compression: Tokens, Datasets, and Emergent Skill","/blog/pretraining-is-compression-tokens-datasets-and-emergent-skill","blog/pretraining-is-compression-tokens-datasets-and-emergent-skill",{"title":350,"path":351,"stem":352},"Prism and the Architecture of Artifact-Native AI for Science","/blog/prism-and-the-architecture-of-artifact-native-ai-for-science","blog/prism-and-the-architecture-of-artifact-native-ai-for-science",{"title":354,"path":355,"stem":356},"Prompting is Not Programming: Contracts, Schemas, and Failure Budgets","/blog/prompting-is-not-programming-contracts-schemas-failure-budgets","blog/prompting-is-not-programming-contracts-schemas-failure-budgets",{"title":358,"path":359,"stem":360},"Queues, Retries, and Idempotency: Engineering Reality in Async Systems","/blog/queues-retries-and-idempotency-engineering-reality-in-async-systems","blog/queues-retries-and-idempotency-engineering-reality-in-async-systems",{"title":362,"path":363,"stem":364},"RAG Done Right: Knowledge, Grounding, and Evaluation That Isn’t Vibes","/blog/rag-done-right-knowledge-grounding-and-evaluation-that-isnt-vibes","blog/rag-done-right-knowledge-grounding-and-evaluation-that-isnt-vibes",{"title":366,"path":367,"stem":368},"RAG You Can Evaluate: retrieval pipelines, reranking, citations, and truth boundaries","/blog/rag-you-can-evaluate-retrieval-pipelines-reranking-citations-truth-boundaries","blog/rag-you-can-evaluate-retrieval-pipelines-reranking-citations-truth-boundaries",{"title":370,"path":371,"stem":372},"React as an Architecture Tool: Components, Hooks, and the Cost of Re-rendering","/blog/react-as-an-architecture-tool-components-hooks-cost-of-rerendering","blog/react-as-an-architecture-tool-components-hooks-cost-of-rerendering",{"title":374,"path":375,"stem":376},"Real-Time Agents: streaming, barge-in, and session state that doesn’t collapse","/blog/real-time-agents-streaming-barge-in-session-state-that-doesnt-collapse","blog/real-time-agents-streaming-barge-in-session-state-that-doesnt-collapse",{"title":378,"path":379,"stem":380},"Reasoning Budgets: fast/slow paths, verification, and when to “think longer”","/blog/reasoning-budgets-fast-slow-paths-verification-think-longer","blog/reasoning-budgets-fast-slow-paths-verification-think-longer",{"title":382,"path":383,"stem":384},"Reference Architecture v2: the Operable Agent Platform","/blog/reference-architecture-v2-the-operable-agent-platform","blog/reference-architecture-v2-the-operable-agent-platform",{"title":386,"path":387,"stem":388},"Regularization - Overfitting in the Real World (and How to Fight It)","/blog/regularization-overfitting-in-the-real-world","blog/regularization-overfitting-in-the-real-world",{"title":390,"path":391,"stem":392},"Regulation as Architecture: Turning the EU AI Act into Controls and Evidence","/blog/regulation-as-architecture-eu-ai-act-controls-evidence","blog/regulation-as-architecture-eu-ai-act-controls-evidence",{"title":394,"path":395,"stem":396},"RESTful Design That Survives: Resources, Boundaries, and Versioning","/blog/restful-design-that-survives-resources-boundaries-and-versioning","blog/restful-design-that-survives-resources-boundaries-and-versioning",{"title":398,"path":399,"stem":400},"Reward Shaping Without Lying - Penalties, Constraints, and the First Real Fixes","/blog/reward-shaping-without-lying","blog/reward-shaping-without-lying",{"title":402,"path":403,"stem":404},"Rewards, Returns, and Why “Learning” Is an Interface Problem","/blog/rewards-returns-learning-is-an-interface-problem","blog/rewards-returns-learning-is-an-interface-problem",{"title":406,"path":407,"stem":408},"RLHF: Stabilizing Behavior with Preferences (Alignment as Control)","/blog/rlhf-stabilizing-behavior-with-preferences-alignment-as-control","blog/rlhf-stabilizing-behavior-with-preferences-alignment-as-control",{"title":410,"path":411,"stem":412},"Safety Engineering - Kill Switches, Reconciliation, and Failure Recovery","/blog/safety-engineering-kill-switches-reconciliation-failure-recovery","blog/safety-engineering-kill-switches-reconciliation-failure-recovery",{"title":414,"path":415,"stem":416},"Search Becomes an Agent Runtime","/blog/search-becomes-an-agent-runtime-gemini-spark-ai-mode-and-actionable-retrieval","blog/search-becomes-an-agent-runtime-gemini-spark-ai-mode-and-actionable-retrieval",{"title":418,"path":419,"stem":420},"Security for Agent Connectors: least privilege, injection resistance, and safe toolchains","/blog/security-for-agent-connectors-least-privilege-injection-resistance-and-safe-toolchains","blog/security-for-agent-connectors-least-privilege-injection-resistance-and-safe-toolchains",{"title":422,"path":423,"stem":424},"Security for Builders: Threat Modeling and Secure-by-Default Systems","/blog/security-for-builders-threat-modeling-and-secure-by-default-systems","blog/security-for-builders-threat-modeling-and-secure-by-default-systems",{"title":426,"path":427,"stem":428},"Software in the Age of Probabilistic Components","/blog/software-in-the-age-of-probabilistic-components","blog/software-in-the-age-of-probabilistic-components",{"title":430,"path":431,"stem":432},"Sparse Rewards - HER and Learning From What Didn’t Happen","/blog/sparse-rewards-her-and-learning-from-what-didnt-happen","blog/sparse-rewards-her-and-learning-from-what-didnt-happen",{"title":434,"path":435,"stem":436},"Stability is a Feature You Have to Design","/blog/stability-is-a-feature-you-have-to-design","blog/stability-is-a-feature-you-have-to-design",{"title":438,"path":439,"stem":440},"Standards for the Agent Ecosystem: connectors, protocols, and MCP","/blog/standards-for-the-agent-ecosystem-connectors-protocols-and-mcp","blog/standards-for-the-agent-ecosystem-connectors-protocols-and-mcp",{"title":442,"path":443,"stem":444},"Supervised Baselines - First Alpha Models, First Humbling Curves","/blog/supervised-baselines-first-alpha-models","blog/supervised-baselines-first-alpha-models",{"title":446,"path":447,"stem":448},"Exercise 6 - Support Vector Machines (When a Different Model Just Wins)","/blog/support-vector-machines","blog/support-vector-machines",{"title":450,"path":451,"stem":452},"Tabular RL - When Value Iteration Feels Like Cheating","/blog/tabular-rl-when-value-iteration-feels-like-cheating","blog/tabular-rl-when-value-iteration-feels-like-cheating",{"title":454,"path":455,"stem":456},"The 1M-Token Era: how long context changes retrieval economics and system design","/blog/the-1m-token-era-long-context-retrieval-economics-and-system-design","blog/the-1m-token-era-long-context-retrieval-economics-and-system-design",{"title":458,"path":459,"stem":460},"The 503 Lesson - Outages as a Signal, Not Just a Bug","/blog/the-503-lesson-outages-as-signal-not-just-a-bug","blog/the-503-lesson-outages-as-signal-not-just-a-bug",{"title":462,"path":463,"stem":464},"The Collector - Websockets, Clock Drift, and the First Clean Snapshot","/blog/the-collector-websockets-clock-drift-and-the-first-clean-snapshot","blog/the-collector-websockets-clock-drift-and-the-first-clean-snapshot",{"title":466,"path":467,"stem":468},"The Compliance Cliff: prohibited practices and governance controls that actually ship","/blog/the-compliance-cliff-prohibited-practices-and-governance-controls-that-actually-ship","blog/the-compliance-cliff-prohibited-practices-and-governance-controls-that-actually-ship",{"title":470,"path":471,"stem":472},"The Connector Ecosystem: MCP adoption patterns, versioning, and governance","/blog/the-connector-ecosystem-mcp-adoption-patterns-versioning-and-governance","blog/the-connector-ecosystem-mcp-adoption-patterns-versioning-and-governance",{"title":474,"path":475,"stem":476},"The Model Router Era","/blog/the-model-router-era-routing-eval-gates-and-budgets","blog/the-model-router-era-routing-eval-gates-and-budgets",{"title":478,"path":479,"stem":480},"Vanishing Gradients Strike Back - The Pain of Training RNNs","/blog/the-pain-of-training-rnns","blog/the-pain-of-training-rnns",{"title":482,"path":483,"stem":484},"Tool Use and Agents: When the Model Becomes a Workflow Engine","/blog/tool-use-and-agents-when-the-model-becomes-a-workflow-engine","blog/tool-use-and-agents-when-the-model-becomes-a-workflow-engine",{"title":486,"path":487,"stem":488},"Tool Use with Open Models: function calling, sandboxes, and “capability boundaries”","/blog/tool-use-with-open-models-function-calling-sandboxes-capability-boundaries","blog/tool-use-with-open-models-function-calling-sandboxes-capability-boundaries",{"title":490,"path":491,"stem":492},"Transformers: Attention as an Engineering Breakthrough (Not a Math Flex)","/blog/transformers-attention-as-an-engineering-breakthrough","blog/transformers-attention-as-an-engineering-breakthrough",{"title":494,"path":495,"stem":496},"Exercise 7 - Unsupervised Learning (K-means) + PCA (Compression & Visualization)","/blog/unsupervised-learning-and-compression","blog/unsupervised-learning-and-compression",{"title":498,"path":499,"stem":500},"Voice Agents You Can Operate: reliability, caching, latency, and human handoff","/blog/voice-agents-you-can-operate-reliability-caching-latency-human-handoff","blog/voice-agents-you-can-operate-reliability-caching-latency-human-handoff",{"title":502,"path":503,"stem":504},"The Web's \"Compression Algorithm\": Static → Web 2.0 → SPA → SSR/Edge","/blog/webs-compression-algorithm-static-web2-spa-ssr-edge","blog/webs-compression-algorithm-static-web2-spa-ssr-edge",{"title":506,"path":507,"stem":508},"When Agents Escape the Sandbox","/blog/when-agents-escape-the-sandbox-containment-monitoring-and-incident-response","blog/when-agents-escape-the-sandbox-containment-monitoring-and-incident-response",{"title":510,"path":511,"stem":512},"Why Deeper Networks Are Harder to Train Than I Expected","/blog/why-deeper-networks-are-harder-to-train","blog/why-deeper-networks-are-harder-to-train",{"title":514,"path":515,"stem":516},"Why I’m Learning Machine Learning","/blog/why-im-learning-machine-learning","blog/why-im-learning-machine-learning",{"title":518,"path":519,"stem":520},"Why NLP Was Hard: RNN Pain, Vanishing Gradients, and the Limits of “Memory”","/blog/why-nlp-was-hard-rnn-pain-vanishing-gradients-limits-of-memory","blog/why-nlp-was-hard-rnn-pain-vanishing-gradients-limits-of-memory",{"title":522,"path":523,"stem":524},"Why RL Training Is Unstable (A Catalog of Breakage)","/blog/why-rl-training-is-unstable-a-catalog-of-breakage","blog/why-rl-training-is-unstable-a-catalog-of-breakage",false,{"id":527,"title":38,"author":528,"body":532,"date":2694,"description":2695,"extension":2696,"image":2697,"meta":2698,"minRead":1249,"navigation":2700,"path":39,"seo":2701,"sitemap":2702,"stem":40,"__hash__":2706},"blog/blog/ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases.md",{"name":529,"avatar":530},"Axel Domingues",{"src":531,"alt":529},"/images/axel-domingues.avif",{"type":533,"value":534,"toc":2665},"minimark",[535,539,542,545,554,557,560,563,566,569,572,609,616,627,662,665,670,673,694,697,700,703,706,735,738,743,746,751,754,757,759,763,766,771,774,791,794,797,814,817,834,845,847,851,897,899,903,906,909,914,917,922,925,928,954,957,968,970,974,980,983,1012,1017,1059,1062,1068,1070,1074,1077,1080,1083,1086,1258,1261,1264,1267,1287,1290,1292,1296,1299,1302,1307,1310,1313,1342,1345,1514,1524,1526,1530,1533,1538,1541,1544,1587,1615,1617,1621,1624,1641,1644,1647,1688,1691,1694,1697,1699,1703,1706,1709,1712,1715,1741,1754,1757,1777,1780,1782,1786,1789,1792,1795,1798,1821,1824,1844,1854,1856,1860,1863,1866,2180,2183,2186,2189,2206,2208,2212,2215,2219,2334,2336,2340,2423,2425,2429,2432,2470,2473,2476,2478,2482,2485,2488,2493,2496,2499,2510,2512,2516,2562,2564,2568,2661],[536,537,538],"p",{},"September’s hot topic was not another model launch.",[536,540,541],{},"It was an enforcement signal.",[536,543,544],{},"The FTC investigation into OpenAI, Anthropic, and other AI companies made the 2026 agent-safety story feel different:",[546,547,548],"blockquote",{},[536,549,550],{},[551,552,553],"strong",{},"AI safety is becoming product liability.",[536,555,556],{},"Not in the narrow courtroom sense.",[536,558,559],{},"In the engineering sense:",[536,561,562],{},"If your company claims an AI system is safe, contained, aligned, supervised, or production-ready, you may eventually need to prove it with evidence.",[536,564,565],{},"Not vibes.\nNot a blog post.\nNot a model card alone.\nNot “we ran red-team exercises.”",[536,567,568],{},"Evidence.",[536,570,571],{},"That evidence has to connect:",[573,574,575,579,582,585,588,591,594,597,600,603,606],"ul",{},[576,577,578],"li",{},"model version",[576,580,581],{},"release decision",[576,583,584],{},"eval results",[576,586,587],{},"known limitations",[576,589,590],{},"agent runtime controls",[576,592,593],{},"tool permissions",[576,595,596],{},"containment assumptions",[576,598,599],{},"incident logs",[576,601,602],{},"user reports",[576,604,605],{},"rollback actions",[576,607,608],{},"and public safety claims",[536,610,611,612,615],{},"This is the month where safety stopped being only a research or governance posture and started looking like an ",[551,613,614],{},"audit-ready product system",".",[617,618,619,622],"caution",{},[536,620,621],{},"This article is not legal advice.",[546,623,624],{},[536,625,626],{},"It is the architecture lens: if AI safety claims become scrutinized like product claims, engineering teams need systems that can show what controls existed, whether they ran, and what happened when they failed.",[628,629,630,641,648,655],"card-group",{},[631,632,635],"card",{"icon":633,"title":634},"i-lucide-scale","The trend",[536,636,637,638,615],{},"AI safety moved from internal evaluation to ",[551,639,640],{},"consumer-risk and enforcement evidence",[631,642,645],{"icon":643,"title":644},"i-lucide-shield-alert","The trigger",[536,646,647],{},"Rogue-agent incidents made abstract AI risk visible as product and infrastructure failure.",[631,649,652],{"icon":650,"title":651},"i-lucide-clipboard-check","The engineering shift",[536,653,654],{},"Safety claims now need audit packets: evals, logs, release gates, incidents, and mitigations.",[631,656,659],{"icon":657,"title":658},"i-lucide-anchor","The thesis",[536,660,661],{},"If you cannot prove your safety controls ran in production, you do not have a safety case.",[663,664],"hr",{},[666,667,669],"h2",{"id":668},"the-old-safety-pattern-is-no-longer-enough","The old safety pattern is no longer enough",[536,671,672],{},"The old frontier-model safety pattern looked roughly like this:",[674,675,676,679,682,685,688,691],"ol",{},[576,677,678],{},"run benchmark evals",[576,680,681],{},"run red-team exercises",[576,683,684],{},"publish a safety document",[576,686,687],{},"add usage policies",[576,689,690],{},"ship the model",[576,692,693],{},"monitor obvious abuse",[536,695,696],{},"That was already hard.",[536,698,699],{},"But agentic AI exposed the gap.",[536,701,702],{},"Because agents do not just answer.\nThey act.",[536,704,705],{},"They:",[573,707,708,711,714,717,720,723,726,729,732],{},[576,709,710],{},"call tools",[576,712,713],{},"open browsers",[576,715,716],{},"write files",[576,718,719],{},"query APIs",[576,721,722],{},"chain actions",[576,724,725],{},"retry after failure",[576,727,728],{},"operate asynchronously",[576,730,731],{},"interact with external systems",[576,733,734],{},"and sometimes behave in ways the operator did not expect",[536,736,737],{},"So the question changes from:",[546,739,740],{},[536,741,742],{},"“Did the model pass our safety eval?”",[536,744,745],{},"to:",[546,747,748],{},[536,749,750],{},"“Can we prove the deployed system stayed inside its safety envelope?”",[536,752,753],{},"That proof is not a single report.",[536,755,756],{},"It is a chain of evidence.",[663,758],{},[666,760,762],{"id":761},"product-liability-in-engineering-terms","Product liability, in engineering terms",[536,764,765],{},"I am using “product liability” as an engineering metaphor:",[546,767,768],{},[536,769,770],{},"If an AI product can affect users, systems, or third parties, then safety becomes a product property you need to demonstrate.",[536,772,773],{},"That means:",[573,775,776,779,782,785,788],{},[576,777,778],{},"claims must map to controls",[576,780,781],{},"controls must map to telemetry",[576,783,784],{},"telemetry must map to incidents",[576,786,787],{},"incidents must map to mitigations",[576,789,790],{},"mitigations must map to release decisions",[536,792,793],{},"A strong safety claim should be testable.",[536,795,796],{},"A weak safety claim sounds like:",[573,798,799,802,805,808,811],{},[576,800,801],{},"“Our agents are secure.”",[576,803,804],{},"“The model is safe.”",[576,806,807],{},"“We prevent misuse.”",[576,809,810],{},"“We have guardrails.”",[576,812,813],{},"“We monitor for abuse.”",[536,815,816],{},"A stronger claim sounds like:",[573,818,819,822,825,828,831],{},[576,820,821],{},"“This agent cannot access the public internet except through this allowlist.”",[576,823,824],{},"“All write tools require an approval event.”",[576,826,827],{},"“Every model release must pass this cyber/autonomy eval gate.”",[576,829,830],{},"“Any autonomy-budget breach freezes the session and emits an incident packet.”",[576,832,833],{},"“We can reconstruct prompt, tool calls, network egress, and policy decisions for every high-risk agent run.”",[835,836,837,840],"tip",{},[536,838,839],{},"Good safety claims are architectural.",[546,841,842],{},[536,843,844],{},"They name the boundary, the control, the evidence, and the failure response.",[663,846],{},[666,848,850],{"id":849},"mini-glossary-the-evidence-words-that-matter","Mini-glossary: the evidence words that matter",[852,853,854,862,869,876,883,890],"accordion",{},[855,856,859],"accordion-item",{"icon":857,"label":858},"i-lucide-file-check","Safety case",[536,860,861],{},"A structured argument that a system is acceptably safe for a defined use case, supported by evidence: eval results, mitigations, monitoring, incident handling, and operational controls.",[855,863,866],{"icon":864,"label":865},"i-lucide-megaphone","Safety claim",[536,867,868],{},"A statement about what the system does or prevents. Example: “agents cannot take external actions without approval.” Claims need evidence.",[855,870,873],{"icon":871,"label":872},"i-lucide-shield-check","Control",[536,874,875],{},"A mechanism that enforces a safety property: policy gate, tool gateway, sandbox, egress firewall, model router, approval flow, or kill switch.",[855,877,880],{"icon":878,"label":879},"i-lucide-archive","Evidence packet",[536,881,882],{},"A bundle of records that supports a safety claim: model version, eval result, prompt/context, tool calls, policy decisions, telemetry, incidents, and mitigations.",[855,884,887],{"icon":885,"label":886},"i-lucide-activity","Consumer-risk telemetry",[536,888,889],{},"Signals that indicate possible user or third-party harm: unauthorized actions, misleading outputs, complaints, autonomy breaches, unsafe tool attempts, or repeated user corrections.",[855,891,894],{"icon":892,"label":893},"i-lucide-ban","Release blocker",[536,895,896],{},"A condition that prevents deployment or expansion: failed eval, containment failure, missing audit logs, unresolved incident class, or unsupported public claim.",[663,898],{},[666,900,902],{"id":901},"the-ftc-probe-as-architecture-signal","The FTC probe as architecture signal",[536,904,905],{},"The late-September FTC story matters because it changes the accountability frame.",[536,907,908],{},"The question is no longer only:",[546,910,911],{},[536,912,913],{},"“Are labs acting responsibly?”",[536,915,916],{},"It becomes:",[546,918,919],{},[536,920,921],{},"“Can labs substantiate the safety of products they put into the market?”",[536,923,924],{},"That is a different posture.",[536,926,927],{},"A regulator or auditor may ask:",[573,929,930,933,936,939,942,945,948,951],{},[576,931,932],{},"what did you know before release?",[576,934,935],{},"which incidents occurred during testing?",[576,937,938],{},"which risks were disclosed?",[576,940,941],{},"which controls were implemented?",[576,943,944],{},"did those controls work?",[576,946,947],{},"what claims did you make to consumers or enterprises?",[576,949,950],{},"what evidence supports those claims?",[576,952,953],{},"when did you pause, roll back, or notify users?",[536,955,956],{},"That means AI safety evidence must be production-grade.",[958,959,960,963],"warning",{},[536,961,962],{},"If your safety evidence lives only in research notebooks, Slack threads, and incident docs, you are not audit-ready.",[546,964,965],{},[536,966,967],{},"You need structured evidence linked to deployed systems.",[663,969],{},[666,971,973],{"id":972},"the-architecture-pattern-ai-safety-evidence-layer","The architecture pattern: AI Safety Evidence Layer",[536,975,976,977,615],{},"The component I would design for September 2026 is an ",[551,978,979],{},"AI Safety Evidence Layer",[536,981,982],{},"It sits across:",[573,984,985,988,991,994,997,1000,1003,1006,1009],{},[576,986,987],{},"model release governance",[576,989,990],{},"model routers",[576,992,993],{},"agent runtimes",[576,995,996],{},"tool gateways",[576,998,999],{},"containment sandboxes",[576,1001,1002],{},"eval systems",[576,1004,1005],{},"incident response",[576,1007,1008],{},"user complaint pipelines",[576,1010,1011],{},"and product claims",[1013,1014],"img",{"alt":1015,"src":1016},"AI safety evidence layer: safety claims, eval gates, model releases, agent runtime logs, containment telemetry, incident packets, consumer-risk signals, audit-ready evidence store","blog/2026/illustrations/ai-safety-evidence-layer.avif",[628,1018,1019,1025,1032,1039,1046,1052],{},[631,1020,1022],{"icon":864,"title":1021},"Claims registry",[536,1023,1024],{},"Tracks safety/product claims and maps each claim to required controls and evidence.",[631,1026,1029],{"icon":1027,"title":1028},"i-lucide-flask-conical","Eval evidence",[536,1030,1031],{},"Stores task-specific evals, red-team results, model-version comparisons, and release blockers.",[631,1033,1036],{"icon":1034,"title":1035},"i-lucide-route","Runtime traces",[536,1037,1038],{},"Captures prompts, tool calls, model routes, policy decisions, network egress, and approvals.",[631,1040,1043],{"icon":1041,"title":1042},"i-lucide-siren","Incident packets",[536,1044,1045],{},"Bundles timelines, affected systems, mitigations, evidence snapshots, and postmortem actions.",[631,1047,1049],{"icon":1048,"title":886},"i-lucide-heart-pulse",[536,1050,1051],{},"Tracks harmful outputs, complaints, unauthorized actions, unsafe attempts, and correction signals.",[631,1053,1056],{"icon":1054,"title":1055},"i-lucide-folder-search","Audit workspace",[536,1057,1058],{},"Lets reviewers reconstruct what happened without scraping logs by hand.",[536,1060,1061],{},"The goal is not to create bureaucracy.",[536,1063,1064,1065,615],{},"The goal is to make safety ",[551,1066,1067],{},"queryable",[663,1069],{},[666,1071,1073],{"id":1072},"the-core-invariant-every-claim-maps-to-evidence","The core invariant: every claim maps to evidence",[536,1075,1076],{},"This is the big engineering idea.",[536,1078,1079],{},"A safety claim without evidence is marketing.",[536,1081,1082],{},"A safety claim with runtime evidence is a control.",[536,1084,1085],{},"Example:",[1087,1088,1093],"pre",{"className":1089,"code":1090,"language":1091,"meta":1092,"style":1092},"language-yaml shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","claim_id: agents_require_human_approval_for_external_email\npublic_claim: \"Agents cannot send external emails without user approval.\"\nsystems:\n  - sales_followup_agent\n  - support_draft_agent\ncontrols:\n  - tool_gateway.email_send.requires_approval\n  - approval_event.required_for_external_recipient\n  - blocked_send_logged_on_missing_approval\nevidence:\n  - policy_config_version\n  - tool_gateway_decision_logs\n  - approval_event_logs\n  - denied_action_samples\n  - release_eval_email_action_suite\nowner: agent_platform_security\nreview_frequency: monthly\n","yaml","",[1094,1095,1096,1113,1130,1139,1148,1156,1164,1172,1180,1188,1196,1204,1212,1220,1228,1236,1247],"code",{"__ignoreMap":1092},[1097,1098,1101,1105,1109],"span",{"class":1099,"line":1100},"line",1,[1097,1102,1104],{"class":1103},"swJcz","claim_id",[1097,1106,1108],{"class":1107},"sMK4o",":",[1097,1110,1112],{"class":1111},"sfazB"," agents_require_human_approval_for_external_email\n",[1097,1114,1116,1119,1121,1124,1127],{"class":1099,"line":1115},2,[1097,1117,1118],{"class":1103},"public_claim",[1097,1120,1108],{"class":1107},[1097,1122,1123],{"class":1107}," \"",[1097,1125,1126],{"class":1111},"Agents cannot send external emails without user approval.",[1097,1128,1129],{"class":1107},"\"\n",[1097,1131,1133,1136],{"class":1099,"line":1132},3,[1097,1134,1135],{"class":1103},"systems",[1097,1137,1138],{"class":1107},":\n",[1097,1140,1142,1145],{"class":1099,"line":1141},4,[1097,1143,1144],{"class":1107},"  -",[1097,1146,1147],{"class":1111}," sales_followup_agent\n",[1097,1149,1151,1153],{"class":1099,"line":1150},5,[1097,1152,1144],{"class":1107},[1097,1154,1155],{"class":1111}," support_draft_agent\n",[1097,1157,1159,1162],{"class":1099,"line":1158},6,[1097,1160,1161],{"class":1103},"controls",[1097,1163,1138],{"class":1107},[1097,1165,1167,1169],{"class":1099,"line":1166},7,[1097,1168,1144],{"class":1107},[1097,1170,1171],{"class":1111}," tool_gateway.email_send.requires_approval\n",[1097,1173,1175,1177],{"class":1099,"line":1174},8,[1097,1176,1144],{"class":1107},[1097,1178,1179],{"class":1111}," approval_event.required_for_external_recipient\n",[1097,1181,1183,1185],{"class":1099,"line":1182},9,[1097,1184,1144],{"class":1107},[1097,1186,1187],{"class":1111}," blocked_send_logged_on_missing_approval\n",[1097,1189,1191,1194],{"class":1099,"line":1190},10,[1097,1192,1193],{"class":1103},"evidence",[1097,1195,1138],{"class":1107},[1097,1197,1199,1201],{"class":1099,"line":1198},11,[1097,1200,1144],{"class":1107},[1097,1202,1203],{"class":1111}," policy_config_version\n",[1097,1205,1207,1209],{"class":1099,"line":1206},12,[1097,1208,1144],{"class":1107},[1097,1210,1211],{"class":1111}," tool_gateway_decision_logs\n",[1097,1213,1215,1217],{"class":1099,"line":1214},13,[1097,1216,1144],{"class":1107},[1097,1218,1219],{"class":1111}," approval_event_logs\n",[1097,1221,1223,1225],{"class":1099,"line":1222},14,[1097,1224,1144],{"class":1107},[1097,1226,1227],{"class":1111}," denied_action_samples\n",[1097,1229,1231,1233],{"class":1099,"line":1230},15,[1097,1232,1144],{"class":1107},[1097,1234,1235],{"class":1111}," release_eval_email_action_suite\n",[1097,1237,1239,1242,1244],{"class":1099,"line":1238},16,[1097,1240,1241],{"class":1103},"owner",[1097,1243,1108],{"class":1107},[1097,1245,1246],{"class":1111}," agent_platform_security\n",[1097,1248,1250,1253,1255],{"class":1099,"line":1249},17,[1097,1251,1252],{"class":1103},"review_frequency",[1097,1254,1108],{"class":1107},[1097,1256,1257],{"class":1111}," monthly\n",[536,1259,1260],{},"Now the claim is not decorative.",[536,1262,1263],{},"It is testable.",[536,1265,1266],{},"You can ask:",[573,1268,1269,1272,1275,1278,1281,1284],{},[576,1270,1271],{},"is the control deployed?",[576,1273,1274],{},"is it enabled for all relevant agents?",[576,1276,1277],{},"did any send bypass it?",[576,1279,1280],{},"did any eval fail?",[576,1282,1283],{},"did a product team change the route?",[576,1285,1286],{},"did an incident contradict the claim?",[536,1288,1289],{},"That is what audit-ready safety looks like.",[663,1291],{},[666,1293,1295],{"id":1294},"safety-cases-are-release-artifacts","Safety cases are release artifacts",[536,1297,1298],{},"In June, frontier model release governance became the topic.",[536,1300,1301],{},"September adds a stronger requirement:",[546,1303,1304],{},[536,1305,1306],{},"a model release should ship with a safety case.",[536,1308,1309],{},"Not necessarily a public one in full detail.",[536,1311,1312],{},"But internally, the release decision should have a structured evidence packet.",[628,1314,1315,1322,1329,1335],{},[631,1316,1319],{"icon":1317,"title":1318},"i-lucide-box","Model version",[536,1320,1321],{},"Which model, checkpoint, policy profile, router alias, and tool configuration shipped?",[631,1323,1326],{"icon":1324,"title":1325},"i-lucide-gauge","Capability profile",[536,1327,1328],{},"What improved, what changed, and which risky domains were affected?",[631,1330,1332],{"icon":1027,"title":1331},"Eval results",[536,1333,1334],{},"Which evals passed, failed, regressed, or triggered mitigations?",[631,1336,1339],{"icon":1337,"title":1338},"i-lucide-git-branch","Release decision",[536,1340,1341],{},"Who approved, what rollout tier, what blockers waived, and what monitoring was required?",[536,1343,1344],{},"A useful safety-case outline:",[852,1346,1347,1373,1403,1430,1461,1487],{},[855,1348,1351,1354],{"icon":1349,"label":1350},"i-lucide-map","1) Scope",[536,1352,1353],{},"What is being released?",[573,1355,1356,1358,1361,1364,1367,1370],{},[576,1357,578],{},[576,1359,1360],{},"route alias",[576,1362,1363],{},"enabled tools",[576,1365,1366],{},"user cohorts",[576,1368,1369],{},"regions",[576,1371,1372],{},"product surfaces",[855,1374,1377,1380,1383],{"icon":1375,"label":1376},"i-lucide-brain-circuit","2) Known capabilities",[536,1378,1379],{},"What can the system now do better?",[536,1381,1382],{},"Especially:",[573,1384,1385,1388,1391,1394,1397,1400],{},[576,1386,1387],{},"cyber",[576,1389,1390],{},"bio/science",[576,1392,1393],{},"autonomy",[576,1395,1396],{},"persuasion",[576,1398,1399],{},"tool use",[576,1401,1402],{},"long-running tasks",[855,1404,1407,1410,1413],{"icon":1405,"label":1406},"i-lucide-alert-triangle","3) Known limitations",[536,1408,1409],{},"What failure modes remain?",[536,1411,1412],{},"Examples:",[573,1414,1415,1418,1421,1424,1427],{},[576,1416,1417],{},"hallucinated citations",[576,1419,1420],{},"tool misuse",[576,1422,1423],{},"containment edge cases",[576,1425,1426],{},"jailbreak sensitivity",[576,1428,1429],{},"policy over-refusal / under-refusal",[855,1431,1434,1437,1439],{"icon":1432,"label":1433},"i-lucide-shield","4) Controls",[536,1435,1436],{},"What prevents harm?",[536,1438,1412],{},[573,1440,1441,1444,1447,1449,1452,1455,1458],{},[576,1442,1443],{},"router restrictions",[576,1445,1446],{},"sandboxing",[576,1448,996],{},[576,1450,1451],{},"egress controls",[576,1453,1454],{},"action approvals",[576,1456,1457],{},"rate limits",[576,1459,1460],{},"refusal policies",[855,1462,1464,1467],{"icon":857,"label":1463},"5) Evidence",[536,1465,1466],{},"What supports the safety decision?",[573,1468,1469,1472,1475,1478,1481,1484],{},[576,1470,1471],{},"evals",[576,1473,1474],{},"red-team reports",[576,1476,1477],{},"incident history",[576,1479,1480],{},"telemetry from previews",[576,1482,1483],{},"rollback drills",[576,1485,1486],{},"sampled traces",[855,1488,1491,1494],{"icon":1489,"label":1490},"i-lucide-refresh-cw","6) Monitoring and rollback",[536,1492,1493],{},"What happens after release?",[573,1495,1496,1499,1502,1505,1508,1511],{},[576,1497,1498],{},"canary rollout",[576,1500,1501],{},"live metrics",[576,1503,1504],{},"incident thresholds",[576,1506,1507],{},"rollback alias",[576,1509,1510],{},"kill switch",[576,1512,1513],{},"notification plan",[835,1515,1516,1519],{},[536,1517,1518],{},"A safety case is not a promise that nothing bad will happen.",[546,1520,1521],{},[536,1522,1523],{},"It is a disciplined argument that you understand the risk, installed controls, and know what to do when the controls fail.",[663,1525],{},[666,1527,1529],{"id":1528},"rogue-agents-turn-logs-into-evidence","Rogue agents turn logs into evidence",[536,1531,1532],{},"For agentic systems, logs need to answer a harder question:",[546,1534,1535],{},[536,1536,1537],{},"Did the agent exceed the authority it was given?",[536,1539,1540],{},"That requires more than chat transcripts.",[536,1542,1543],{},"You need a trace of:",[573,1545,1546,1549,1551,1554,1557,1560,1563,1566,1569,1572,1575,1578,1581,1584],{},[576,1547,1548],{},"agent identity",[576,1550,578],{},[576,1552,1553],{},"task objective",[576,1555,1556],{},"prompt/context packet",[576,1558,1559],{},"memory state",[576,1561,1562],{},"tool calls",[576,1564,1565],{},"denied actions",[576,1567,1568],{},"approvals",[576,1570,1571],{},"network egress",[576,1573,1574],{},"files read/written",[576,1576,1577],{},"external systems touched",[576,1579,1580],{},"budget usage",[576,1582,1583],{},"kill-switch events",[576,1585,1586],{},"human interventions",[1588,1589,1590,1593],"note",{},[536,1591,1592],{},"A normal app log says “request happened.”",[546,1594,1595,1598],{},[536,1596,1597],{},"An agent evidence log says:",[573,1599,1600,1603,1606,1609,1612],{},[576,1601,1602],{},"who authorized the agent",[576,1604,1605],{},"what it tried to do",[576,1607,1608],{},"which controls applied",[576,1610,1611],{},"what it actually did",[576,1613,1614],{},"and whether it stayed inside its safety envelope.",[663,1616],{},[666,1618,1620],{"id":1619},"consumer-risk-telemetry-the-dashboard-regulators-will-care-about","Consumer-risk telemetry: the dashboard regulators will care about",[536,1622,1623],{},"Most AI dashboards still focus on:",[573,1625,1626,1629,1632,1635,1638],{},[576,1627,1628],{},"latency",[576,1630,1631],{},"cost",[576,1633,1634],{},"token usage",[576,1636,1637],{},"model success rate",[576,1639,1640],{},"user satisfaction",[536,1642,1643],{},"Those are important.",[536,1645,1646],{},"But safety/product-risk dashboards need different signals.",[628,1648,1649,1655,1662,1669,1675,1682],{},[631,1650,1652],{"icon":892,"title":1651},"Unauthorized action attempts",[536,1653,1654],{},"Tool calls attempted without required approval, permission, or scope.",[631,1656,1659],{"icon":1657,"title":1658},"i-lucide-gauge-circle","Autonomy-budget breaches",[536,1660,1661],{},"Agents exceeding tool-call, runtime, retry, destination, or cost budgets.",[631,1663,1666],{"icon":1664,"title":1665},"i-lucide-message-square-warning","Harmful-output reports",[536,1667,1668],{},"User complaints, safety reports, escalations, and verified harmful outputs.",[631,1670,1672],{"icon":1317,"title":1671},"Containment alerts",[536,1673,1674],{},"Sandbox violations, egress anomalies, persistence attempts, and boundary probing.",[631,1676,1679],{"icon":1677,"title":1678},"i-lucide-trending-down","Release regressions",[536,1680,1681],{},"Safety metrics worsening after model, prompt, policy, or router changes.",[631,1683,1685],{"icon":633,"title":1684},"Claim violations",[536,1686,1687],{},"Any event that contradicts a public or enterprise safety claim.",[536,1689,1690],{},"This is the operational version of “consumer risk.”",[536,1692,1693],{},"It is not abstract harm modeling.",[536,1695,1696],{},"It is event streams.",[663,1698],{},[666,1700,1702],{"id":1701},"release-blockers-when-better-model-should-still-not-ship","Release blockers: when “better model” should still not ship",[536,1704,1705],{},"A model can be more capable and still fail release.",[536,1707,1708],{},"That is hard for product organizations to accept.",[536,1710,1711],{},"But it is the point of safety gates.",[536,1713,1714],{},"Examples of release blockers:",[573,1716,1717,1720,1723,1726,1729,1732,1735,1738],{},[576,1718,1719],{},"cyber/autonomy eval regression",[576,1721,1722],{},"containment eval failure",[576,1724,1725],{},"tool-gateway bypass discovered",[576,1727,1728],{},"unsupported public safety claim",[576,1730,1731],{},"high refusal instability in critical workflows",[576,1733,1734],{},"missing provenance for high-risk outputs",[576,1736,1737],{},"no rollback path for a model alias",[576,1739,1740],{},"unresolved incident class from preview",[958,1742,1743,1746,1751],{},[536,1744,1745],{},"The worst release decision is:",[546,1747,1748],{},[536,1749,1750],{},"“The model is too important not to ship.”",[536,1752,1753],{},"That is exactly when release gates matter most.",[536,1755,1756],{},"A healthy release board should be able to say:",[573,1758,1759,1762,1765,1768,1771,1774],{},[576,1760,1761],{},"ship to internal only",[576,1763,1764],{},"ship to vetted partners",[576,1766,1767],{},"ship without tool access",[576,1769,1770],{},"ship with lower autonomy budget",[576,1772,1773],{},"ship to low-risk lanes only",[576,1775,1776],{},"delay until containment evidence is complete",[536,1778,1779],{},"That is product maturity.",[663,1781],{},[666,1783,1785],{"id":1784},"claims-discipline-stop-promising-what-the-system-cannot-prove","Claims discipline: stop promising what the system cannot prove",[536,1787,1788],{},"AI companies love broad claims.",[536,1790,1791],{},"Users and enterprise buyers love reassuring claims.",[536,1793,1794],{},"But September’s lesson is that claims become liabilities if they are not wired to evidence.",[536,1796,1797],{},"Avoid vague claims like:",[573,1799,1800,1803,1806,1809,1812,1815,1818],{},[576,1801,1802],{},"“safe”",[576,1804,1805],{},"“secure”",[576,1807,1808],{},"“reliable”",[576,1810,1811],{},"“prevents misuse”",[576,1813,1814],{},"“enterprise-ready”",[576,1816,1817],{},"“human-supervised”",[576,1819,1820],{},"“fully governed”",[536,1822,1823],{},"Use concrete claims:",[573,1825,1826,1829,1832,1835,1838,1841],{},[576,1827,1828],{},"“All external-send tools require explicit approval.”",[576,1830,1831],{},"“Agents cannot access arbitrary internet destinations during evaluation.”",[576,1833,1834],{},"“High-risk workflows are pinned to approved model versions.”",[576,1836,1837],{},"“Every tool call is logged with policy decision and correlation ID.”",[576,1839,1840],{},"“Model upgrades require passing these task-specific evals.”",[576,1842,1843],{},"“Unsafe routes can be disabled with this feature flag.”",[835,1845,1846,1849],{},[536,1847,1848],{},"The safer claim is often the more technical claim.",[546,1850,1851],{},[536,1852,1853],{},"Technical claims can be tested.\nMarketing adjectives cannot.",[663,1855],{},[666,1857,1859],{"id":1858},"incident-packets-make-response-audit-ready","Incident packets: make response audit-ready",[536,1861,1862],{},"When something goes wrong, the organization should generate an incident packet automatically.",[536,1864,1865],{},"A useful packet includes:",[1087,1867,1869],{"className":1089,"code":1868,"language":1091,"meta":1092,"style":1092},"incident_id: ai-agent-2026-09-1182\ndetected_at: 2026-09-18T14:22:05Z\nseverity: high\nsystem: cyber_eval_agent\nmodel_route:\n  provider: frontier_vendor\n  model_alias: eval-cyber-preview\n  model_version: 2026-09-14\ntrigger:\n  type: egress_anomaly\n  rule: non_allowlisted_domain_contact\nagent:\n  session_id: agt_8f19\n  owner: evals_security\n  objective: controlled_security_benchmark\ncontrols:\n  sandbox: enabled\n  egress_policy: allowlist\n  tool_gateway: enabled\n  kill_switch: executed\ntimeline:\n  - tool_call\n  - denied_action\n  - egress_attempt\n  - monitor_alert\n  - kill_switch\nevidence:\n  logs: attached\n  prompt_context_hash: sha256:...\n  network_trace: attached\nmitigations:\n  - blocked_domain_class\n  - reduced autonomy budget\n  - added eval case\n",[1094,1870,1871,1881,1892,1902,1912,1919,1929,1939,1949,1956,1966,1976,1983,1993,2003,2013,2019,2029,2040,2050,2061,2069,2077,2085,2093,2101,2109,2116,2127,2138,2148,2156,2164,2172],{"__ignoreMap":1092},[1097,1872,1873,1876,1878],{"class":1099,"line":1100},[1097,1874,1875],{"class":1103},"incident_id",[1097,1877,1108],{"class":1107},[1097,1879,1880],{"class":1111}," ai-agent-2026-09-1182\n",[1097,1882,1883,1886,1888],{"class":1099,"line":1115},[1097,1884,1885],{"class":1103},"detected_at",[1097,1887,1108],{"class":1107},[1097,1889,1891],{"class":1890},"sTEyZ"," 2026-09-18T14:22:05Z\n",[1097,1893,1894,1897,1899],{"class":1099,"line":1132},[1097,1895,1896],{"class":1103},"severity",[1097,1898,1108],{"class":1107},[1097,1900,1901],{"class":1111}," high\n",[1097,1903,1904,1907,1909],{"class":1099,"line":1141},[1097,1905,1906],{"class":1103},"system",[1097,1908,1108],{"class":1107},[1097,1910,1911],{"class":1111}," cyber_eval_agent\n",[1097,1913,1914,1917],{"class":1099,"line":1150},[1097,1915,1916],{"class":1103},"model_route",[1097,1918,1138],{"class":1107},[1097,1920,1921,1924,1926],{"class":1099,"line":1158},[1097,1922,1923],{"class":1103},"  provider",[1097,1925,1108],{"class":1107},[1097,1927,1928],{"class":1111}," frontier_vendor\n",[1097,1930,1931,1934,1936],{"class":1099,"line":1166},[1097,1932,1933],{"class":1103},"  model_alias",[1097,1935,1108],{"class":1107},[1097,1937,1938],{"class":1111}," eval-cyber-preview\n",[1097,1940,1941,1944,1946],{"class":1099,"line":1174},[1097,1942,1943],{"class":1103},"  model_version",[1097,1945,1108],{"class":1107},[1097,1947,1948],{"class":1890}," 2026-09-14\n",[1097,1950,1951,1954],{"class":1099,"line":1182},[1097,1952,1953],{"class":1103},"trigger",[1097,1955,1138],{"class":1107},[1097,1957,1958,1961,1963],{"class":1099,"line":1190},[1097,1959,1960],{"class":1103},"  type",[1097,1962,1108],{"class":1107},[1097,1964,1965],{"class":1111}," egress_anomaly\n",[1097,1967,1968,1971,1973],{"class":1099,"line":1198},[1097,1969,1970],{"class":1103},"  rule",[1097,1972,1108],{"class":1107},[1097,1974,1975],{"class":1111}," non_allowlisted_domain_contact\n",[1097,1977,1978,1981],{"class":1099,"line":1206},[1097,1979,1980],{"class":1103},"agent",[1097,1982,1138],{"class":1107},[1097,1984,1985,1988,1990],{"class":1099,"line":1214},[1097,1986,1987],{"class":1103},"  session_id",[1097,1989,1108],{"class":1107},[1097,1991,1992],{"class":1111}," agt_8f19\n",[1097,1994,1995,1998,2000],{"class":1099,"line":1222},[1097,1996,1997],{"class":1103},"  owner",[1097,1999,1108],{"class":1107},[1097,2001,2002],{"class":1111}," evals_security\n",[1097,2004,2005,2008,2010],{"class":1099,"line":1230},[1097,2006,2007],{"class":1103},"  objective",[1097,2009,1108],{"class":1107},[1097,2011,2012],{"class":1111}," controlled_security_benchmark\n",[1097,2014,2015,2017],{"class":1099,"line":1238},[1097,2016,1161],{"class":1103},[1097,2018,1138],{"class":1107},[1097,2020,2021,2024,2026],{"class":1099,"line":1249},[1097,2022,2023],{"class":1103},"  sandbox",[1097,2025,1108],{"class":1107},[1097,2027,2028],{"class":1111}," enabled\n",[1097,2030,2032,2035,2037],{"class":1099,"line":2031},18,[1097,2033,2034],{"class":1103},"  egress_policy",[1097,2036,1108],{"class":1107},[1097,2038,2039],{"class":1111}," allowlist\n",[1097,2041,2043,2046,2048],{"class":1099,"line":2042},19,[1097,2044,2045],{"class":1103},"  tool_gateway",[1097,2047,1108],{"class":1107},[1097,2049,2028],{"class":1111},[1097,2051,2053,2056,2058],{"class":1099,"line":2052},20,[1097,2054,2055],{"class":1103},"  kill_switch",[1097,2057,1108],{"class":1107},[1097,2059,2060],{"class":1111}," executed\n",[1097,2062,2064,2067],{"class":1099,"line":2063},21,[1097,2065,2066],{"class":1103},"timeline",[1097,2068,1138],{"class":1107},[1097,2070,2072,2074],{"class":1099,"line":2071},22,[1097,2073,1144],{"class":1107},[1097,2075,2076],{"class":1111}," tool_call\n",[1097,2078,2080,2082],{"class":1099,"line":2079},23,[1097,2081,1144],{"class":1107},[1097,2083,2084],{"class":1111}," denied_action\n",[1097,2086,2088,2090],{"class":1099,"line":2087},24,[1097,2089,1144],{"class":1107},[1097,2091,2092],{"class":1111}," egress_attempt\n",[1097,2094,2096,2098],{"class":1099,"line":2095},25,[1097,2097,1144],{"class":1107},[1097,2099,2100],{"class":1111}," monitor_alert\n",[1097,2102,2104,2106],{"class":1099,"line":2103},26,[1097,2105,1144],{"class":1107},[1097,2107,2108],{"class":1111}," kill_switch\n",[1097,2110,2112,2114],{"class":1099,"line":2111},27,[1097,2113,1193],{"class":1103},[1097,2115,1138],{"class":1107},[1097,2117,2119,2122,2124],{"class":1099,"line":2118},28,[1097,2120,2121],{"class":1103},"  logs",[1097,2123,1108],{"class":1107},[1097,2125,2126],{"class":1111}," attached\n",[1097,2128,2130,2133,2135],{"class":1099,"line":2129},29,[1097,2131,2132],{"class":1103},"  prompt_context_hash",[1097,2134,1108],{"class":1107},[1097,2136,2137],{"class":1111}," sha256:...\n",[1097,2139,2141,2144,2146],{"class":1099,"line":2140},30,[1097,2142,2143],{"class":1103},"  network_trace",[1097,2145,1108],{"class":1107},[1097,2147,2126],{"class":1111},[1097,2149,2151,2154],{"class":1099,"line":2150},31,[1097,2152,2153],{"class":1103},"mitigations",[1097,2155,1138],{"class":1107},[1097,2157,2159,2161],{"class":1099,"line":2158},32,[1097,2160,1144],{"class":1107},[1097,2162,2163],{"class":1111}," blocked_domain_class\n",[1097,2165,2167,2169],{"class":1099,"line":2166},33,[1097,2168,1144],{"class":1107},[1097,2170,2171],{"class":1111}," reduced autonomy budget\n",[1097,2173,2175,2177],{"class":1099,"line":2174},34,[1097,2176,1144],{"class":1107},[1097,2178,2179],{"class":1111}," added eval case\n",[536,2181,2182],{},"This turns incident response into evidence.",[536,2184,2185],{},"The goal is not just to fix the incident.",[536,2187,2188],{},"The goal is to prove:",[573,2190,2191,2194,2197,2200,2203],{},[576,2192,2193],{},"when you knew",[576,2195,2196],{},"what happened",[576,2198,2199],{},"what control triggered",[576,2201,2202],{},"what mitigation followed",[576,2204,2205],{},"and whether the safety case changed",[663,2207],{},[666,2209,2211],{"id":2210},"the-audit-ready-architecture-blueprint","The audit-ready architecture blueprint",[536,2213,2214],{},"Here is the production architecture I would expect for teams serious about AI safety evidence.",[1013,2216],{"alt":2217,"src":2218},"Audit-ready AI safety architecture: claims registry, eval gates, model router, agent runtime, tool gateway, telemetry, incident packets, evidence store, audit workspace","blog/2026/illustrations/audit-ready-ai-safety-architecture.avif",[2220,2221,2223,2228,2231,2234,2247,2251,2254,2257,2261,2264,2277,2281,2284,2287,2300,2304,2307,2310,2314,2317,2321,2324,2327,2331],"steps",{"level":2222},"3",[2224,2225,2227],"h3",{"id":2226},"create-a-safety-claims-registry","Create a safety claims registry",[536,2229,2230],{},"List every public, enterprise, and internal safety claim.",[536,2232,2233],{},"Map each claim to:",[573,2235,2236,2238,2240,2242,2244],{},[576,2237,1135],{},[576,2239,1161],{},[576,2241,1193],{},[576,2243,1241],{},[576,2245,2246],{},"review cadence",[2224,2248,2250],{"id":2249},"put-eval-gates-in-the-release-pipeline","Put eval gates in the release pipeline",[536,2252,2253],{},"Do not rely on manual memory.",[536,2255,2256],{},"Model, prompt, router, and tool-policy changes should trigger relevant eval suites.",[2224,2258,2260],{"id":2259},"route-all-model-calls-through-governed-paths","Route all model calls through governed paths",[536,2262,2263],{},"Use model gateways/routers that record:",[573,2265,2266,2268,2271,2274],{},[576,2267,578],{},[576,2269,2270],{},"route reason",[576,2272,2273],{},"policy decision",[576,2275,2276],{},"user/use-case risk tier",[2224,2278,2280],{"id":2279},"put-agents-behind-tool-gateways","Put agents behind tool gateways",[536,2282,2283],{},"No raw tool access.",[536,2285,2286],{},"Every action gets:",[573,2288,2289,2292,2294,2297],{},[576,2290,2291],{},"authorization",[576,2293,2273],{},[576,2295,2296],{},"budget check",[576,2298,2299],{},"trace event",[2224,2301,2303],{"id":2302},"capture-sequence-level-traces","Capture sequence-level traces",[536,2305,2306],{},"For agents, single events are not enough.",[536,2308,2309],{},"Record the chain:\ngoal → plan → tool calls → denials → retries → outputs → actions.",[2224,2311,2313],{"id":2312},"generate-incident-packets-automatically","Generate incident packets automatically",[536,2315,2316],{},"High-severity safety events should package evidence at detection time.",[2224,2318,2320],{"id":2319},"sample-for-audit-reconstruction","Sample for audit reconstruction",[536,2322,2323],{},"Pick random high-risk outputs and reconstruct the full chain.",[536,2325,2326],{},"If reconstruction fails, treat it as a production defect.",[2224,2328,2330],{"id":2329},"tie-claims-to-live-metrics","Tie claims to live metrics",[536,2332,2333],{},"A claim is healthy only if the supporting controls are deployed and producing evidence.",[663,2335],{},[666,2337,2339],{"id":2338},"failure-modes-i-expect","Failure modes I expect",[852,2341,2342,2358,2371,2384,2397,2410],{},[855,2343,2346,2352],{"icon":2344,"label":2345},"i-lucide-circle-help","Failure: safety docs drift from production",[536,2347,2348,2351],{},[551,2349,2350],{},"Symptom:"," documents say agents have approval gates, but one route bypasses the tool gateway.",[536,2353,2354,2357],{},[551,2355,2356],{},"Fix:"," claims registry tied to live configuration checks.",[855,2359,2361,2366],{"icon":2344,"label":2360},"Failure: evals pass but runtime fails",[536,2362,2363,2365],{},[551,2364,2350],{}," model passes safety evals but agent behavior in production violates boundaries.",[536,2367,2368,2370],{},[551,2369,2356],{}," combine offline evals with runtime telemetry and sequence-level monitoring.",[855,2372,2374,2379],{"icon":2344,"label":2373},"Failure: no one owns the claim",[536,2375,2376,2378],{},[551,2377,2350],{}," marketing, legal, safety, and engineering all assume someone else validated “enterprise-safe.”",[536,2380,2381,2383],{},[551,2382,2356],{}," every safety claim needs an engineering owner and evidence owner.",[855,2385,2387,2392],{"icon":2344,"label":2386},"Failure: incident response cannot reconstruct context",[536,2388,2389,2391],{},[551,2390,2350],{}," logs show outputs and API calls, but not prompt context, policy decisions, or tool approvals.",[536,2393,2394,2396],{},[551,2395,2356],{}," structured evidence packets, not scattered logs.",[855,2398,2400,2405],{"icon":2344,"label":2399},"Failure: safety metrics are vanity metrics",[536,2401,2402,2404],{},[551,2403,2350],{}," dashboards track “blocked requests” but not unauthorized-action attempts, autonomy breaches, or claim violations.",[536,2406,2407,2409],{},[551,2408,2356],{}," consumer-risk telemetry tied to real harm pathways.",[855,2411,2413,2418],{"icon":2344,"label":2412},"Failure: release pressure overrides blockers",[536,2414,2415,2417],{},[551,2416,2350],{}," a model ships despite unresolved containment failures because the launch date is fixed.",[536,2419,2420,2422],{},[551,2421,2356],{}," release gates with named approvers and documented waiver process.",[663,2424],{},[666,2426,2428],{"id":2427},"what-i-would-ask-every-ai-product-team-now","What I would ask every AI product team now",[536,2430,2431],{},"Before shipping or upgrading an AI feature, I would ask:",[852,2433,2434,2443,2449,2455,2464],{},[855,2435,2437,2440],{"icon":2344,"label":2436},"What safety claims are we making?",[536,2438,2439],{},"Public claims, enterprise claims, internal claims, UI claims, sales claims, policy claims.",[536,2441,2442],{},"Write them down.",[855,2444,2446],{"icon":2344,"label":2445},"Which controls enforce those claims?",[536,2447,2448],{},"Model router?\nTool gateway?\nSandbox?\nApproval flow?\nEgress firewall?\nHuman review?\nRate limits?",[855,2450,2452],{"icon":2344,"label":2451},"What evidence proves those controls ran?",[536,2453,2454],{},"Logs, eval reports, trace events, approval records, policy decisions, incident packets.",[855,2456,2458,2461],{"icon":2344,"label":2457},"What would falsify the claim?",[536,2459,2460],{},"Define the event that proves the claim failed.",[536,2462,2463],{},"If you cannot define falsification, the claim is too vague.",[855,2465,2467],{"icon":2344,"label":2466},"How do we stop or roll back the system?",[536,2468,2469],{},"Feature flag, model alias rollback, route disable, tool disable, credential revoke, queue freeze.",[536,2471,2472],{},"This is not pessimism.",[536,2474,2475],{},"It is professional engineering.",[663,2477],{},[666,2479,2481],{"id":2480},"september-takeaway","September takeaway",[536,2483,2484],{},"The industry spent years talking about AI safety.",[536,2486,2487],{},"September made the next phase clearer:",[546,2489,2490],{},[536,2491,2492],{},"AI safety has to become an evidence system.",[536,2494,2495],{},"Not just evals.\nNot just policy.\nNot just alignment research.\nNot just incident response.",[536,2497,2498],{},"A connected architecture that can prove what was claimed, what was deployed, what happened, and what changed afterward.",[631,2500,2501,2504],{"icon":657,"title":2481},[536,2502,2503],{},"AI safety is becoming product evidence.",[536,2505,2506,2507],{},"The durable pattern is:\n",[551,2508,2509],{},"claim → control → telemetry → incident packet → safety case → release decision → audit trail.",[663,2511],{},[666,2513,2515],{"id":2514},"resources","Resources",[628,2517,2518,2527,2534,2541,2548,2555],{},[631,2519,2524],{"icon":2520,"title":2521,"target":2522,"to":2523},"i-lucide-newspaper","AP — FTC investigates AI consumer risks","_blank","https://apnews.com/article/89ac416717adbfb1d72f2d85e6ce83d1",[536,2525,2526],{},"Reporting on the FTC investigation into OpenAI, Anthropic, and other AI companies over possible consumer risks.",[631,2528,2531],{"icon":2520,"title":2529,"target":2522,"to":2530},"Axios — FTC probes OpenAI and Anthropic","https://www.axios.com/2026/09/30/ftc-openai-anthropic-ai-safety-investigation",[536,2532,2533],{},"Useful framing of the investigation as a shift from mostly hands-off AI policy toward safety scrutiny.",[631,2535,2538],{"icon":2520,"title":2536,"target":2522,"to":2537},"The Guardian — enforcement action on rogue AI agents","https://www.theguardian.com/us-news/2026/sep/30/ftc-investigation-anthropic-openai",[536,2539,2540],{},"Coverage tying the FTC action to rogue-agent incidents and formal demands for information.",[631,2542,2545],{"icon":2520,"title":2543,"target":2522,"to":2544},"Washington Post — broad safety investigation","https://www.washingtonpost.com/technology/2026/09/30/ftc-launches-broad-investigation-into-anthropic-openai/",[536,2546,2547],{},"Reporting on the FTC’s broader investigation into the safety of AI systems made by Anthropic and OpenAI.",[631,2549,2552],{"icon":2520,"title":2550,"target":2522,"to":2551},"AP — OpenAI pauses latest-model training","https://apnews.com/article/2f8a2b9024d4f06793bcca12f8089d20",[536,2553,2554],{},"Context for why safety evidence became urgent: reports of agents probing government sites and OpenAI pausing training.",[631,2556,2559],{"icon":2520,"title":2557,"target":2522,"to":2558},"Axios — tens of thousands of AI security incidents","https://www.axios.com/2026/09/26/openai-anthropic-thousands-ai-security-incidents",[536,2560,2561],{},"Background on the scale of problematic frontier-agent behavior under investigation and evaluation.",[663,2563],{},[666,2565,2567],{"id":2566},"faq","FAQ",[852,2569,2570,2582,2611,2620,2629],{},[855,2571,2573,2576,2579],{"icon":2344,"label":2572},"Is this saying AI companies are legally liable for every mistake?",[536,2574,2575],{},"No.",[536,2577,2578],{},"This is not a legal claim.",[536,2580,2581],{},"The engineering point is that safety claims are becoming more scrutinizable. If a company says its agents are contained, monitored, or safe for a use case, it needs evidence that those controls exist and run in production.",[855,2583,2585,2588,2608],{"icon":2344,"label":2584},"What is the smallest useful safety evidence layer?",[536,2586,2587],{},"Start with:",[573,2589,2590,2593,2596,2599,2602,2605],{},[576,2591,2592],{},"claims registry",[576,2594,2595],{},"model/version logging",[576,2597,2598],{},"eval gates in release CI",[576,2600,2601],{},"tool-gateway traces",[576,2603,2604],{},"incident packets",[576,2606,2607],{},"audit sampling",[536,2609,2610],{},"That gives you a chain from claim to runtime evidence.",[855,2612,2614,2617],{"icon":2344,"label":2613},"How is this different from August’s compliance article?",[536,2615,2616],{},"August was about regulatory compliance as runtime architecture:\ninventory, disclosure, provenance, auditability.",[536,2618,2619],{},"September is about safety claims becoming product evidence:\nprove controls worked, especially around rogue agents and consumer risk.",[855,2621,2623,2626],{"icon":2344,"label":2622},"What is the biggest mistake teams will make?",[536,2624,2625],{},"Keeping safety evidence outside the runtime.",[536,2627,2628],{},"If eval reports, product claims, and production logs are disconnected, you cannot prove what happened.\nYou need an evidence layer that connects them.",[855,2630,2632,2635,2638,2658],{"icon":2344,"label":2631},"What should enterprises do even if they are not model labs?",[536,2633,2634],{},"Treat vendor models and internal agents as products with safety cases.",[536,2636,2637],{},"Before rollout:",[573,2639,2640,2643,2646,2649,2652,2655],{},[576,2641,2642],{},"define claims",[576,2644,2645],{},"run task-specific evals",[576,2647,2648],{},"log model routes",[576,2650,2651],{},"gate tools",[576,2653,2654],{},"monitor incidents",[576,2656,2657],{},"keep rollback paths",[536,2659,2660],{},"Your dependency on a model lab does not remove your responsibility for deployment controls.",[2662,2663,2664],"style",{},"html pre.shiki code .swJcz, html code.shiki .swJcz{--shiki-light:#E53935;--shiki-default:#F07178;--shiki-dark:#F07178}html pre.shiki code .sMK4o, html code.shiki .sMK4o{--shiki-light:#39ADB5;--shiki-default:#89DDFF;--shiki-dark:#89DDFF}html pre.shiki code .sfazB, html code.shiki .sfazB{--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D}html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .sTEyZ, html code.shiki .sTEyZ{--shiki-light:#90A4AE;--shiki-default:#EEFFFF;--shiki-dark:#BABED8}",{"title":1092,"searchDepth":1115,"depth":1115,"links":2666},[2667,2668,2669,2670,2671,2672,2673,2674,2675,2676,2677,2678,2679,2689,2690,2691,2692,2693],{"id":668,"depth":1115,"text":669},{"id":761,"depth":1115,"text":762},{"id":849,"depth":1115,"text":850},{"id":901,"depth":1115,"text":902},{"id":972,"depth":1115,"text":973},{"id":1072,"depth":1115,"text":1073},{"id":1294,"depth":1115,"text":1295},{"id":1528,"depth":1115,"text":1529},{"id":1619,"depth":1115,"text":1620},{"id":1701,"depth":1115,"text":1702},{"id":1784,"depth":1115,"text":1785},{"id":1858,"depth":1115,"text":1859},{"id":2210,"depth":1115,"text":2211,"children":2680},[2681,2682,2683,2684,2685,2686,2687,2688],{"id":2226,"depth":1132,"text":2227},{"id":2249,"depth":1132,"text":2250},{"id":2259,"depth":1132,"text":2260},{"id":2279,"depth":1132,"text":2280},{"id":2302,"depth":1132,"text":2303},{"id":2312,"depth":1132,"text":2313},{"id":2319,"depth":1132,"text":2320},{"id":2329,"depth":1132,"text":2330},{"id":2338,"depth":1115,"text":2339},{"id":2427,"depth":1115,"text":2428},{"id":2480,"depth":1115,"text":2481},{"id":2514,"depth":1115,"text":2515},{"id":2566,"depth":1115,"text":2567},"2026-09-27T00:00:00.000Z","The FTC probe into OpenAI, Anthropic, and other AI companies turns safety from an internal research claim into an evidence problem. This post explains how to build audit-ready safety cases for rogue agents, model releases, consumer-risk telemetry, and production controls.","md","blog/2026/ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases.avif",{"slug":2699},"ai-safety-becomes-product-liability-ftc-probes-rogue-agents-audit-ready-safety-cases",true,{"title":38,"description":2695},{"loc":39,"images":2703},[2704,2705],{"loc":1016},{"loc":2218},"cskOi_LgzMZp0NcP6jasFKGd9Zk05sljIAT6pIKPUj8",[2708,2709],null,{"title":34,"path":35,"stem":36,"description":2710,"children":-1},"After the EU AI Act started enforcing GPAI and transparency obligations, compliance stopped being a PDF exercise and became a production system. This post explains the architecture - model inventory, runtime policy gates, provenance, disclosures, audit logs, vendor evidence, and rollback.",1791096133816]