AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements

arXiv CS Tuesday 09 June 2026, 04:00 UTC By Zhiwei Liu, Yueru He, Qing Ou, Tianlei Zhu, Xiaorui Guo, Xueqing Peng, Sophia Ananiadou 1 min read

Key Points

Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in financial analysis and surface-level factual error detection, yet their ability to identify fraudulent financial misinformation in audited corporate reporting remains underexplored. Existing financial and audit benchmarks mainly focus on factual verification, numerical reasoning, rule compliance, or audit workflows, but rarely evaluate misleading disclosure narratives or management explanations that...

arXiv:2606.08345v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in financial analysis and surface-level factual error detection, yet their ability to identify fraudulent financial misinformation in audited corporate reporting remains underexplored. Existing financial and audit benchmarks mainly focus on factual verification, numerical reasoning, rule compliance, or audit workflows, but rarely evaluate misleading disclosure narratives or management explanations that obscure the true drivers of reported performance. We introduce AuditFraudBench, an enforcement-grounded benchmark constructed from authentic company filings and regulatory materials, including original and restated 10-K and 10-Q filings, structured financial statements, MD&A disclosures, and SEC Accounting and Auditing Enforcement Releases (AAERs). AuditFraudBench contains three tasks: Profit Source Attribution, Misleading Narrative Detection, and Fraud Pattern Classification, which evaluate whether models can identify the true source of reported performance, detect misleading disclosure framing, and classify misconduct mechanisms into known manipulation patterns. We evaluate GPT, DeepSeek, and Qwen series LLMs on the benchmark. Results show that both proprietary and open models still struggle to jointly reason over financial figures, disclosure framing, restatement evidence, and enforcement-grounded fraud mechanisms. AuditFraudBench provides a challenging testbed for audit-relevant, evidence-grounded evaluation of LLMs in financial reporting.

AuditFraudBench (ORG) SEC Accounting and Auditing Enforcement Releases (ORG) Fraud Pattern Classification (ORG) GPT (ORG) Qwen (PERSON)

Originally published by arXiv CS Read original →

AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements

Related Stories

Hot potato! Get rid of these imposters while you c...

School uniform charity plans fundraising week

M&S announces £30million of price cuts on dozens of food lines - see full list

Trump Says Citi Is the Top M&A Adviser, But It's Not