Patronus AI Launches Lynx: State-of-the-Art Open Source Hallucination Detection Model
New hallucination evaluation benchmark shows that the new model is more accurate at catching hallucinations than GPT-4o, GPT-4-Turbo, Claude-3 and industry solutions
Hallucinations occur when LLMs generate responses that are coherent but do not align with factual reality or the input context, undermining their practical utility across various applications. While traditional proprietary LLMs, like GPT-4, have become used to detect these inconsistencies in recent times ('LLM-as-a-judge'), there are concerns over their reliability, scalability, and cost.
Lynx represents a breakthrough in the field by enabling real-time hallucination detection without the need for manual annotation. Patronus AI also open sourced HaluBench, a new benchmark sourced from real-world domains, to assess faithfulness in LLM responses comprehensively.
"Since the release of ChatGPT in
Lynx is the first model that beats GPT-4 on hallucination tasks. Lynx (70B) achieved the highest accuracy at detecting hallucinations, compared to all other LLMs used as judges, making it the largest and most powerful open source hallucination model to date. It outperformed OpenAI's GPT models and Anthropic's Claude 3 models at a fraction of the size.
Lynx and HaluBench also support real world domains like Finance and Medicine, which previous datasets and models did not include, making it more applicable to real world problems.
Results:
- In medical answers (PubMedQA), Lynx (70B) was 8.3% more accurate than GPT-4o at detecting medical inaccuracies.
- Lynx (8B) outperformed GPT-3.5 by 24.5% on HaluBench, and beat Claude-3-Sonnet and Claude-3-Haiku by 8.6% and 18.4% respectively, showing strong capabilities in a smaller model.
- Both Lynx (8B) and Lynx (70B) achieve significantly increased accuracy compared to open source model baselines, with Lynx (8B) showing gains of 13.3% over Llama-3-8B-Instruct from supervised finetuning.
- Lynx (70B) outperformed GPT-3.5 by an average of 29.0% across all tasks.
Lynx and HaluBench are now publicly available on Hugging Face, the open source AI platform.
About Patronus AI
Patronus AI is the first automated evaluation and security platform that helps companies use large language models (LLMs) safely. For more information, visit https://www.patronus.ai/ or reach out to [email protected].
View original content:https://www.prnewswire.com/news-releases/patronus-ai-launches-lynx-state-of-the-art-open-source-hallucination-detection-model-302194659.html
SOURCE Patronus AI
Serious News for Serious Traders! Try StreetInsider.com Premium Free!
You May Also Be Interested In
- Trip.com Group Releases 2025 Sustainability Report, Announces New Global Paid Paternity Leave Policy
- Crypto Lags Wall Street’s Rally as Next Crypto To Explode Buyers Use FINAL30 on AlphaPepe Before August 1
- UPDATED: DARPA Selects IonQ to Produce Next-Generation Atomic Clocks
Create E-mail Alert Related Categories
PRNewswire, Press ReleasesSign up for StreetInsider Free!
Receive full access to all new and archived articles, unlimited portfolio tracking, e-mail alerts, custom newswires and RSS feeds - and more!



Tweet
Share