PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

PACT is a benchmark for evaluating the compliance of Large Language Model (LLM) agents in enterprise settings under pressure. It measures how well LLMs follow rules in sensitive contexts, such as hiring, healthcare, and finance, and highlights compliance risks in LLM assistants.

RSS Score 0 9/17/2026, 4:00:00 AM Original Source
Save an API key to vote.