Sep 17, 2026
Tencent's Hy4 Preview evaluated across our benchmark suite
We evaluated Tencentโs open-weight Hy4 Preview across our benchmark suite.
-
#21 of 58 on the Vals Index (55.40%) at $1.28 per test, the #4 open-weight model behind DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3.
-
Best open-weight model weโve tested on Code Migration (47.43%, #7 of 61, at $3.41 per test versus $24.91 for GLM 5.3), Public Benefits Bench (68.61%, #4 of 36, behind only three Anthropic models) and SRE Bench (2.29%, #7 of 10).
-
Strong on reasoning and legal work: 75.00% on ProofBench v1.1 (#8 of 32), 59.33% on IOI (#10 of 28), 45.19% on Legal Research Bench (#10 of 61) and 9.17% on Harveyโs Legal Agent Benchmark (#13 of 62). Weaker on Terminal-Bench 2.1 (55.06%, #42 of 66), MedCode (43.25%, #37 of 93) and MedScribe (83.60%, #35 of 95).
-
Cost per test is below the board median on nearly every benchmark, skewing higher on benchmarks that involve files. Latency is medium-high for a model of its size and price (measured at low concurrency): closest to peers on math and science, slowest on routine knowledge work such as MedCode, Legal Research and Public Benefits Bench.
Hy4 Preview is a reasoning model with a 1M-token context window and 64k max output tokens, priced at $0.83 per million input tokens and $2.50 per million output tokens. We accessed it through OpenRouter.
Congrats to the Tencent team on the release!