l8r Twats Library

@jun_song

Post

If LLM-as-a-verifier works as well as my last post suggests, this is going to be massive for local AI.

Like the dev pointed out, pairing a large model with a small model lets you run verification on the cheap.

This means if you combine the GLM-5.3 API with Deepseek-V4-Flash running on a DGX Spark, you get a huge jump in performance without driving up costs.

If you have a 4xDGX Spark, you could multi-batch both GLM and Deepseek on a single node and beat Fable entirely offline without touching an API.

Welcome to the era of local AI.

Quoted post by Jacky Kwok (@jackyk02) @jun\_song Of course! In our paper (https://llm-as-a-verifier.com), we actually used Gemini 2.5 Flash to verify GPT-5.5 trajectories, and it got SOTA on Terminal-Bench 2.0 at the time. So local V4 Flash verifying GLM 5.3 trajectories should definitely work 👀

Open quoted post on X

Explanation

What it says: Jun Song argues that LLM-as-a-verifier could make local AI much more capable: let a stronger model generate candidate reasoning/trajectories, while a cheaper smaller model verifies or selects among them. He suggests GLM-5.3 via API + DeepSeek-V4-Flash locally on a DGX Spark, or a 4× DGX Spark setup running both models locally, could deliver a large performance jump.

Context: The quoted researcher says their paper used Gemini 2.5 Flash to verify GPT-5.5 trajectories, reaching then-SOTA on Terminal-Bench 2.0. They therefore expect a local V4 Flash model verifying GLM-5.3 trajectories to work too. The precise verification algorithm, compute requirements, benchmark numbers, and what “beat Fable” means are missing here.

Why it matters: If the claim generalizes, inference becomes a system-design problem rather than simply “run the biggest model you can.” A relatively small local verifier could cheaply extract more performance from expensive model-generated trajectories, potentially shifting advanced agentic workloads toward local hardware. The key thing to verify is whether the gains survive across benchmarks and whether verifier compute + multiple sampled trajectories actually beat spending the same compute directly on the stronger model.