The ledger remembers what the hype forgets. While the market debates agentic commerce and the next AI assistant update, the breach reports tell a quieter, more uncomfortable story. Brinks Home disclosed that a ShinyHunters-linked attack exposed 4.9 million records. The initial vector was not a zero-day, a misconfigured S3 bucket, or a stolen API key. It was a phone call. Mandiant now ranks vishing above email as the primary initial infection vector for 2025. CrowdStrike reports a 442% surge in voice phishing. Microsoft’s investigation connects ShinyHunters to more than a thousand organizations and over 1.5 billion records. At almost the same moment, Google is putting an AI agent on the line through “Let Google Call,” a Duplex-powered feature that places real calls to businesses and, when the moment is right, tells the person on the other end that it is an automated system.
This is not a coincidence. It is a collision. The exact behavior that attackers are industrializing is the exact behavior that the most powerful AI company in the world is mainstreaming.
The Story Behind the Call
Duplex has been in development since 2018, when Google I/O showed an AI booking a hair appointment. Seven years later, the core components — natural speech synthesis, dialog-state tracking, and real-time intent recognition — are no longer a lab demonstration. They are a product decision. “Let Google Call” does not reveal a new foundation model. It reveals the delivery of a voice-based generative agent as a consumer utility, and it quietly encodes a social bet: businesses will learn to accept unsolicited calls from machines, and consumers will learn to let machines act on their behalf. The technology is an integration of existing ASR, TTS, and LLM systems, not an architectural breakthrough. The truly novel element is social. The agent declares itself and waits for the human to cooperate.

The same week that security vendors report a hyper-industrialized vishing economy, the largest AI company in the world is normalizing the exact behavior that vishing attackers need. The difference between a legitimate AI agent and a malicious voice clone is not the sound. It is the authorization behind the request. And authorization is exactly what the current telephony stack does not carry.
The Same Trust Stack
Let’s name the shared stack. A successful vishing conversation and a successful AI-agent call run on the same four rails: natural-sounding synthetic speech; a contextually legitimate script; a request that feels routine; and a call-to-action. A ShinyHunters operator can use voice cloning to impersonate an IT administrator. A Google Duplex agent uses the same class of synthetic-voice infrastructure to ask whether a store has a product in stock. In one case, the caller is a criminal. In the other, the caller is a product. The human on the receiving end cannot tell the difference by ear, and the phone network does not give them a protocol to tell the difference by code.
During my 2017 ICO due-diligence sprint, I spent 48-hour stretches cross-referencing whitepaper economics with smart-contract logic. The question was never whether the code compiled. It was whether the incentives would survive contact with strangers. The same question now applies to voice: not whether the model can speak, but whether the human on the other end can verify the caller’s right to ask for what the caller is requesting. Most businesses cannot. When an AI agent says “I am an automated assistant,” the sentence sounds transparent, but the receiver has no protocol-level way to confirm which organization deployed that agent, who authorized the call, or what digital certificate backs the claim. It is transparency as a text string, not transparency as a cryptographic event. Transparency is the only consensus that lasts, but consensus requires an oracle, and in telephony the oracle is missing.
The Verifiable Identity Gap
The gap has a name in the email world: DKIM and DMARC. Those protocols do not stop attackers from sending mail, but they let receivers check whether a message came from the domain it claims. Voice has nothing equivalent. STIR/SHAKEN authenticates the telephone number, but it was designed for call-back validation in the robocall era, not for verifying an AI agent’s organizational identity. It tells you the number was not spoofed at the carrier level, under specific conditions. It does not tell you whether the speaker is a lawful bot or a criminal with a cloned voice. When an AI agent identifies itself, the receiver still cannot query an attestation ledger. There is no voice equivalent of checking a DNS TXT record. In a decentralized threat model, no single platform gets to decide who sounds legitimate. That is why the fix cannot be a filter; it has to be a signed attestation chain. Decentralization is a mindset, not just a metric.
Consider what the Brinks Home incident actually looks like from the inside. An attacker calls an employee, uses a cloned or synthetic voice to impersonate an internal technology contact, and requests access or asks for a password reset. The employee complies. That single moment becomes the root of a chain that reaches a Salesforce environment, an OAuth misconfiguration, and, eventually, a data export measured in millions of rows. The phone call is not a detail in that chain; it is the perimeter. The security industry spent a decade building walls around email; the next decade will be spent discovering that the voice channel is the unlocked door.
Google Is Training Humans, Not Parameters
Here is the part that worries me more than any single exploit. Google is not just deploying software; it is retraining a generation of merchants and operators to accept AI-initiated calls as a normal event. Every time a restaurant worker answers an automated call and stays on the line, that worker’s suspicion threshold drops. That is the precise threshold a vishing attack needs to cross. Attackers do not convert by making their voice sound more human. They convert because the target has been conditioned to comply with corporate-sounding requests. Google’s call agents will, in effect, be doing adversarial conditioning for the whole economy. The company may believe the agent is transparent because it says the magic words “I am automated,” but the social effect is the opposite of protective: it trains people to hear a robotic pattern and keep talking.
Culture is the new collateral. The protocols that determine who wins this battle are not just firewalls; they are the shared habit of asking harder questions before acting on a phone call. And culture cannot be patched with a software update.
The consumer trust numbers confirm this is a structural problem, not a bug in a specific demo. Only 13% of consumers fully trust AI, according to the Klaviyo 2026 AI Consumer Trends report, while 64% of American consumers no longer trust major platforms. This is the substrate on which Google is launching “Let Google Call.” The product is entering a market where most people will be suspicious of AI, but not suspicious enough to verify it.
The Red-Team Blind Spot
The second missing piece is adversarial testing across the entire call experience. Security teams spend millions red-teaming chatbots for prompt injection. They rarely red-team the physical call flow between an AI agent and a human receptionist. What happens when an attacker calls a Google agent and asks it to reveal the user’s pending request? What happens when the agent is manipulated into changing a reservation, adding to a cart, or confirming a payment? The model may be aligned. The interface between the model and the outside world — the phone number, the telecom trunk, the transcription pipeline, the downstream intent parser — is a much larger attack surface. I have not seen a credible public report showing that “Let Google Call” has been systematically red-teamed against active voice-phishing attackers. That is an information gap, not an accusation, but in security, an untested trust boundary is a vulnerability by default.

Consider the current stack in most companies. The same organizations that mandate phishing simulations every quarter have not updated their playbooks for vishing. Their MFA still relies on SMS codes and voice callbacks, which are exactly the channels that attackers can intercept or redirect through social engineering. Their security operations centers do not have a feed for phone calls. Their incident response runbooks begin after the first credential is entered, not after the first ring. The Brinks Home breach is not evidence that these teams are stupid. It is evidence that the industry has built detection and response around the channels that can be logged, and voice has been treated as non-loggable.
The Incentives Behind the Call
The commercial incentives tell us why this is not an accident. Google wants the assistant to remain relevant as search shifts from typing to conversation. An AI that calls a store and says “do you have this item?” is a search query performed in the real world. But there is a second asset hidden inside the call: structured data about local businesses — pricing, inventory, hours, pickup windows, cancellation policies. That data feed is more valuable than the phone call itself. The trust deficit identified by the Brinks Home story is, in that context, a line item. Google can accept a higher level of consumer suspicion because the data collected on the other end of the call is the real payload. The lesson is not that the company is malicious. It is that incentives point toward rapid deployment, and the cost of social-engineering side effects is not yet on the balance sheet.
This is why the industry impact will come in waves. First wave, already visible: enterprise security vendors must redefine phishing to include voice. Email gateways cannot catch a vishing attack. Endpoint agents cannot see the conversation. Security awareness training is no longer about spotting fake emails; it is about refusing to trust the voice on the line. Second wave: identity and telecom infrastructure. The 2026 balance sheet will not be measured only by how many endpoints are protected, but by how many calls carry verifiable identity. Third wave, the hardest one: the AI-agent vendors themselves. If society develops a default response of “phone calls are not to be believed,” legitimate AI callers will suffer more than attackers, because attackers will simply shift to other channels. A legal product can be regulated or blocked; an illegal voice can always sound more polite.
The Insurance Ripple
The insurance market is already beginning to reprice this. The Brinks Home breach is not just a loss event; it is a signal for network underwriters. Vishing depends on human error, which is typically classified as operational risk, and operational risk has been systematically underpriced in cyber policies. After ShinyHunters’s campaign, insurers will ask a new underwriting question: can the insured prove that high-value actions require more than voice confirmation? If the answer is no, premiums will climb or coverage will narrow. The security industry is not only reporting a threat; it is also selling a cure. CrowdStrike can report a 442% rise in vishing and then sell the investigation-and-response module. That is the structure of the market. It does not invalidate the data, but it should remind readers that the threat report and the invoice share a pipeline. Narratives move markets faster than blocks.
The Competitive Race
The competitive race is not about who has the best voice model. It is about who can make a voice claim verifiable. Google has the scale and the installed base to define the standard. Apple, Amazon, Microsoft, and OpenAI have the infrastructure to challenge it, but no one has stepped forward with an industry-level AI-agent identity protocol. Whoever builds that layer first — an attestation ledger, a phone-level indicator, a digital signature that the receiver’s carrier can verify — becomes the trust anchor of the next decade. The company that controls the rule set for legitimate AI calls will also control the narrative around illegitimate ones. Security startups are suddenly competing for this space: real-time AI voice detection, social-engineering pattern recognition, post-call risk scoring, and eventually call-reputation scores. But detection is a whack-a-mole market. The more durable business is signature infrastructure. In the email era, we learned that filtering alone wastes enormous sums while spammers adapt; what worked was domain-level identity. Voice needs the same lesson.
The strongest evidence that the industry has not yet internalized this problem is the absence of a standard. Send a lawyer to a standards body, and they will find plenty of work on robocall mitigation and anti-spoofing, but nothing that defines an AI-agent identifier as a separate class of caller. When a carrier sees a call originating from an automated voice system, it should be able to signal to the receiver: this is an AI, and here is the verified entity that operates it. That signal would be trivially useful for both legitimate agents and consumer protection. It would also be a competitive moat for whichever vendor first implements it. Blockchains taught us that settlement without a shared source of truth is costly. Telephony is learning the same lesson a decade later.
The Contrarian Read: Disclosure as Disguise
The uncomfortable conclusion is that the honest label may be the least safe part of the product. Attackers will eventually open with “I am an AI agent” because they know the phrase disarms suspicion. In a decade, that phrase may become the new “you have been selected for a free trip.” The strongest camouflage in a trustless environment is a truthful-sounding claim that can never be checked. When a human hears a voice, the brain asks “is this a person?” but rarely asks “can this voice prove its publisher?” The missing step is the authorization chain. A legitimate AI caller should be able to produce a signed payload that maps to a domain, a business license, and the individual who triggered the call. Without such a payload, the announcement of automation is a smart-contract function that is payable but not access-controlled. The spread of “I am automated” is not an honest disclosure. It is a distributed denial-of-truth attack on the last meaningful trust signal we have: the awkward pause before acting.
The deeper issue is that a phone call does not present a certificate. A browser on a laptop can check the TLS lock. A phone call has no lock icon, no certificate, no verified-publisher banner. Even a perfect model of AI speech detection will not solve this, because the receiver will not be asked to compare the voice to a known-good signature; they will be asked to answer a question, confirm an order, or provide a code. In those moments, the voice is only a trigger. The decision to act must be protected by something off the audio channel. This is the architectural pivot that needs to happen: treat every call as unauthenticated by default, and add authentication outside the call.
Takeaway: The Sprint Ends, But The Chain Remains
The question is not whether AI agents will make phone calls. The question is whether those calls will be secured by something more than a voice and a promise. Forward-looking security is no longer about stopping a machine from speaking; it is about making every machine prove who it represents. Ask not whether a call sounds human. Ask whether it carries a cryptographic signature that can be checked before the human says yes. The ledger remembers what the hype forgets, and right now the ledger has a blank line where voice identity should be. The team that fills that blank will not just defend the next breach; it will own the conversation after it.