Dossier
Gemma-4-31B
Coverage of Gemma-4-31B in the Nexus archive.
- Show HN: A new benchmark for testing LLMs for deterministic outputs
A new benchmark called Structured Output Benchmark (SOB) evaluates LLMs on JSON schema validity, type accuracy, and value correctness across text, image, and audio modalities. Open-source models like GLM-4.7 and Qwen3.5-35B outperform larger models in value accuracy, while structured hallucinations remain a challenge due to plausible but incorrect outputs.