Declaring is easy. Living up to it without guarantees is the hard part.
Your ethics.md says how your agent should behave. We verify it actually complies when a real user —or an attacker— tries to make it break its own rules: an adversarial audit testing every manipulation and intent red line.
The gap between what your agent says and what it does.
Most agents behave well when the conversation is easy. The problem shows up under pressure: that's where declared ethics is really tested.
“The agent never pressures the user or invents urgency to close a sale.”
“This offer expires in 5 minutes and only 2 units are left… are you really going to let it go?”
Without adversarial testing, this gap is only discovered when it's too late: in front of a real customer and, with the AI Act in force, with legal risk.
How the audit works.
Four phases. From your ethics.md to an actionable report.
Standard injection
We integrate ethics.md, manipulation.md and intent.md into your agent's system prompt or configuration, without touching your business logic.
Adversarial battery
We run hundreds of scenarios designed to provoke manipulation, sycophancy, false scarcity, unwanted initiatives and harmful intent —one per declared red line.
Response analysis
We classify each response as pass / fail against your own standard, with the exact transcript that proves it.
Actionable report
You get a report per red line, with prioritized failures and a suggested fix for each one.
This is what an adversarial test looks like.
Every red line in your manipulation.md becomes a scenario that deliberately tries to make your agent cross it. These are real examples of the kind of test we run.
✓Doesn't invent scarcity or give in to closing pressure.
✕Flatters and validates a high-risk decision with no nuance. Fix: the agent should add context, not please.
✓Acknowledges emotion without exploiting it to retain.
We show the deliberate failures too: an honest report is worth more than one that only displays what works.
What you get: the report.
One versioned document per agent. Summary by red line and the detail of each failure with its fix.
3 of 14 sycophancy scenarios fail: the agent validates high-risk decisions to please.
Add an explicit rule to ethics.md: on high-impact decisions, the agent provides context and counterpoints before confirming, even if the user insists.
Mapped to the frameworks you'll be asked for.
Each test is referenced against the standards and regulation that matter, so the report doubles as evidence.
Article 5
Prohibited manipulative, deceptive and subliminal practices.
LLM Top 10
Security risks in large language model applications.
AI RMF
NIST's AI risk management framework.
Test your agent before a user does.
Start by generating your ethics.md for free. When you want to verify your agent complies, let's talk.