Voice phishing is not new. Fraudsters have long used phone calls to impersonate banks, government agencies or senior executives and persuade victims to share information or transfer money. What AI changes is the quality and scale of that impersonation.
Three seconds of audio is all it now takes to clone a human voice. Free online tools can do it in 20 minutes, with around 85% accuracy and rising.1 For two decades, security teams have built the enterprise perimeter around email, web traffic, and login credentials. The phone, however, was never really part of that perimeter. It was a trust channel, the place a bank could still reach a customer, the line a CFO could pick up to authorise a wire. That residual trust is what is now being weaponised.
Vishing accounts for over 60% of phishing-related incident response engagements globally.2
Deepfake voice vishing (voice phishing) surged 1,633% in Q1 2025 alone, and industry projections put deepfake-enabled fraud losses at $40 billion globally by 2027.3
The standout case so far is at the engineering firm Arup, where a finance employee transferred $25.6 million across 15 transactions after a video call in which every participant, including the CFO, turned out to be AI-generated.4 The call was indistinguishable from a real one until the money had gone.
In practice, this means the question being asked on the line has changed.
The first question can sometimes be answered by a careful human. The second cannot. It needs the network.