3.7 Flash reveals sturdy positive factors over 3.6 Flash in coding duties like debugging and problem decision. It additionally achieves larger first-pass code accuracy and has improved efficiency in producing production-ready code as seen in FrontierCode 1.1 Major (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%).
In net growth, 3.7 Flash generates extra purposeful layouts and feature-complete apps in fewer prompts. For UI era, the mannequin reveals excessive design adherence and parity based mostly on a reference enter, whether or not it’s a screenshot, a picture, or a full design system. It outperforms 3.6 Flash on Enviornment.ai’s WebDev Enviornment with an Elo rating of 1588 vs 1538.
For knowledge-dense fields like finance, legislation, and biosciences, 3.7 Flash delivers improved reasoning and accuracy. It considerably outperforms 3.6 Flash on the GDP.pdf benchmark (34.0% vs 22.0%), an eval for testing a mannequin’s skill to course of complicated paperwork. It additionally surpasses 3.6 Flash in AutomationBench, demonstrating it could extra successfully full real-world enterprise workflows (30.4% vs 17.0%).







