← AI security in plain English

AML.T0067

MITRE ATLAS / ATLAS Techniques

Adversaries use prompts that manipulate components of a model's response, such as citations, source links, or formatting, to make the answer appear more trustworthy than it is.

Think of it likeAttaching an official-looking seal and reference number to a letter you wrote yourself.

In plain English

People judge an answer by its trappings. Fabricated citations, invented source links, and confident formatting make a poisoned response read like a well-researched one.