A call comes at an inconvenient hour. Your grandson is on the line, and he is crying. There has been an accident, or an arrest, or a problem at a border, and money has to move inside the next twenty minutes, and please do not tell his mother.

The script is old. Family emergency scams have worked by telephone for decades, and the earlier versions leaned on a poor connection, a name mumbled just enough to pass, and a listener willing to complete the picture themselves. What has changed is the voice. It no longer has to be a stranger’s guess at a stranger’s grandson. It can be built from a recording, and the recording need not be long.

The US Federal Trade Commission put its name to this in March 2023, in a consumer alert by Alvaro Puig titled Scammers use AI to enhance their family emergency schemes. It is short and unhysterical, and its point is that a familiar voice is no longer evidence of who is speaking.

Where the three seconds comes from

Three seconds has a traceable origin. In January 2023, a Microsoft Research team led by Chengyi Wang published a paper on a system called VALL-E. The authors report that a three-second enrolled recording was enough for it to produce speech in the voice of a speaker it had never encountered. They add that the system tends to carry across the emotion and acoustic environment of that sample.

That second detail is easy to over-read. It describes fidelity to the sample: three calm seconds produce calm output, not manufactured panic. Flat affect and a wrong acoustic setting are the artefacts people assume they would catch, and they are the artefacts these systems were built to eliminate.

VALL-E was a research system rather than a product, and the figure describes what that model managed in the authors’ tests. Commercial tools differ in what they ask for and what they return. Read it as a marker of how low the floor had dropped by early 2023, not as a specification.

A birthday video, a voicemail greeting, a few seconds of someone laughing in the background of a reel: supply is not the hard part.

The ear is not a reliable instrument for this

In August 2023, PLOS ONE published work by Kimberly Mai, Sergi Bray, Toby Davies and Lewis Griffin of University College London, titled Warning: Humans cannot reliably detect speech deepfakes. Mai and colleagues played genuine and synthesised audio to 529 people across English and Mandarin and asked them to pick out the fakes. Listeners identified them correctly 73 per cent of the time. Showing them examples beforehand improved matters only slightly.

The finding belongs to that task and that dataset, and it has limits. Those participants had been told they were being tested. They were listening for forgery, with nothing at stake, to voices that were not their own family’s. Someone answering the phone at eleven at night is not performing a detection task, and the paper does not measure whether deep familiarity with a voice helps or hinders.

Seventy-three per cent is probably the generous end of it.

What the FBI’s figures capture, and what they miss

The Internet Crime Complaint Center’s 2025 annual report logged 1,008,597 complaints and $20.9 billion in reported losses. Complainants aged 60 and over accounted for 201,266 of those complaints and $7.7 billion of those losses. Across all age groups, 22,364 complaints carried the report’s AI descriptor, with losses of $893 million.

Distress scams, which the report describes as grandparent scams using voice cloning to imitate a loved one, are a far smaller line item: victims claimed losses of more than $5 million in 2025. That cuts both ways. It is small against the investment and cryptocurrency fraud that dominates elder-fraud losses, and anyone presenting cloned audio as the largest financial threat to older adults is misreading the table. The FBI adds that victims often do not realise AI was involved at all. Plenty of people who hang up feel foolish and tell nobody. Both leave the tagged total as a floor.

The part of the call that has not changed

What has been consistent across the reporting on these calls, well before the audio improved, is the shape of the demand. A deadline. A request for secrecy, usually framed as protecting someone else in the family from worry. And a payment method that cannot be reversed: a wire, gift card numbers and PINs, cryptocurrency, or cash handed to a courier at the door.

The FTC’s guidance on family emergency scams is built around defusing that pressure instead of detecting fakery. In our reading that is the right emphasis. A convincing voice keeps you from asking questions. Everything else in the call exists to stop you making the one phone call that would settle it.

Regulation is running behind the tools

On 8 February 2024 the Federal Communications Commission adopted a declaratory ruling confirming that AI-generated voices fall within the Telephone Consumer Protection Act’s restrictions on artificial or prerecorded voices. The commission’s own news release announced that it had made AI-generated voices in robocalls illegal, which is a stronger headline than the ruling supports. What the ruling did was bring such calls under an existing regime and hand state attorneys general a clearer route to act.

That is a real shift in enforcement posture. A barrier is a different thing, and nothing in the ruling reaches an operator working from another jurisdiction behind a spoofed number.

So the defence is procedural. Hang up, and call back on the number you already hold for that person, not one supplied during the call. Agree a question or a short phrase in advance with the people who might one day ring you in real trouble. Pick something that could not be reconstructed from anything either of you has posted.

It is an odd thing to arrange over Sunday lunch. That is also the only time it can be arranged calmly.