Dossier
structured hallucinations
Coverage of structured hallucinations in the Nexus archive.
- Show HN: A new benchmark for testing LLMs for deterministic outputs
A new benchmark called Structured Output Benchmark (SOB) evaluates LLMs on JSON schema validity, type accuracy, and value correctness across text, image, and audio modalities. Open-source models like GLM-4.7 and Qwen3.5-35B outperform larger models in value accuracy, while structured hallucinations remain a challenge due to plausible but incorrect outputs.