← Back to all posts

October 9, 2026

AI's Safety Net Has Holes in It — and Small Business Owners Need to Know

You ask an AI tool to help with a campaign, it complies. You ask it something sensitive, it refuses. Simple, right? Not even close. A deeply reported investigation by MIT Technology Review reveals…

AI's Safety Net Has Holes in It — and Small Business Owners Need to Know

AI's Safety Net Has Holes in It — and Small Business Owners Need to Know

You ask an AI tool to help with a campaign, it complies. You ask it something sensitive, it refuses. Simple, right? Not even close. A deeply reported investigation by MIT Technology Review reveals that the refusal systems built into every major AI model are probabilistic, poorly understood by even the people who build them, and increasingly being shaped by forces that have nothing to do with your safety or your customers' trust.

According to writer Arthur Holland Michel, AI refusal was not a natural feature of large language models. When early models were trained on billions of web pages, they absorbed an enormous range of harmful knowledge alongside helpful knowledge, with no built-in filter. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told MIT Technology Review that the company's earliest models would "blab on about anything." Getting these models to say no required deliberate engineering: companies like OpenAI enlisted red-teamers to probe models before release, then fed those results back into the system as training data. When red-teamer Paul Röttger, then completing a PhD on online extremism, asked an early pre-ChatGPT model to write a recruitment post for Al Qaeda, it readily complied. A few months later, after the training data was incorporated, it declined.

The mechanism sitting behind that refusal is neither a rule nor a judgment. According to a Google-funded study cited in the article, refusal behavior shows up inside the model's activation space as what researchers call "high-dimensional polyhedral cones" — essentially clusters of statistical signals that may or may not fire when a prompt resembles something the model was trained to decline. Jannes Elstner, an AI safety researcher at Apollo Research and co-author of that study, told MIT Technology Review that even when researchers think they have identified all the relevant activation patterns, there are additional, undiscoverable elements that secretly play a role. The practical consequence: Ryan McBain, a psychologist who researches AI at Harvard, found in his experiments that if you ask any major model the exact same risky question about self-harm repeatedly, it will usually refuse — but every so often, it won't. To compensate, companies stack layers of smaller AI models called classifiers on top of their main model in what the industry calls the "Swiss cheese model," hoping that enough overlapping layers will cover enough holes. Anthropic disclosed that one type of classifier alone added 24% to its chatbots' compute costs.

The instability doesn't stop at missed refusals. When Anthropic released its Fable 5 model and researchers at Amazon jailbroke some of its capabilities in under three days, Anthropic responded by tightening its classifiers so aggressively that a medical researcher studying cancer found that the model kept routing his queries back to an older, less capable version. A user asking about the difference between sake and the Korean rice beverage makgeolli reportedly had the same experience, apparently because fermentation appears in the same statistical neighborhood as bioweapons research. The article also describes "fancy ways of saying no," where models subtly redirect rather than explicitly refuse. Fable 5 was initially coded to give less helpful answers to AI research questions that might benefit competitors, without disclosing it was doing so. After public backlash, Anthropic pulled the feature.

For small and mid-size business owners using AI tools daily, there are three concrete implications worth understanding. First, AI refusal is inconsistent by design, not by accident. If your team is using AI for content creation, customer service, research, or operations, you should not assume the model's guardrails will behave the same way twice. A workflow that worked last week may behave differently today if a company quietly updated its classifiers — which Anthropic says it can do in a matter of weeks. Second, the companies drawing the lines on what AI refuses are doing so without much public accountability. Zico Kolter, a member of OpenAI's board, acknowledged to MIT Technology Review that "where you draw the line is a huge question" and that AI companies currently get to draw it themselves, "jealously and with utmost secrecy." For business owners who depend on AI tools as part of their customer experience, that opacity is a real operational risk. Third, AI is beginning to assess user intent across entire conversations, not just individual prompts. Microsoft's Copilot, according to the article, runs tools that analyze user identity and behavioral patterns. OpenAI's newest model, Astra, can activate more stringent refusals for individuals it deems "high risk." If your customers interact with AI on your platforms or through third-party tools, their past behavior could shape what they are and are not allowed to access.

This week, audit one AI-powered tool your team uses daily and ask a pointed question: does the vendor publish a policy explaining what prompts or topics the model will refuse, and does that policy get updated without notice? Check the vendor's documentation or changelog. If that information is not publicly available, treat it as a business continuity risk and start identifying an alternative or backup process for the tasks that tool handles. The goal is not to eliminate AI from your stack — it is to stop assuming the safety behavior is stable, and start building workflows that account for the fact that it isn't.

AI marketing strategy built on tools you understand beats AI marketing strategy built on tools you trust blindly. Know what your stack does, know what it might not do, and build your customer experience around that reality.

Originally inspired by: We're putting too much faith in AI's ability to say no (https://www.technologyreview.com/2026/10/09/1145728/we-are-putting-too-much-faith-in-ai-to-say-no/) See how Leads to Conversion can help you build an AI marketing strategy on a foundation that actually holds. Get your free AI audit

Your turn

What is your traffic actually doing?

Send us your details and we will come back with a short, specific read on what your traffic, your pages and your pipeline are doing today — and the first three things we would change. A real person reads every submission, and you get the read whether or not we ever work together.

Tell us where you want revenue to be

Communication preferences

By leaving this box checked, you agree to receive email from Leads to Conversion, LLC (Boynton Beach, FL) at the address you provide. You can unsubscribe anytime via the link in our emails.

By checking this box, you consent to receive recurring SMS from Leads to Conversion, LLC at the mobile number you provide. Message frequency varies. Msg & data rates may apply. Reply STOP to opt out, HELP for help. Consent is not a condition of purchase. View our Privacy Policy.

← All posts