Image created with gemini-3.1-flash-image-preview with claude-opus-4.7. Image prompt: A flat matte op-art composition in the style of Julio Le Parc on a clean off-white background, where a single thick ribbon of concentric ROYGBIV bands (violet outside stepping inward to red) loops into the silhouette of a judge’s gavel and sound block, then extends to spell the words ‘AI Inn of Court’ in bold letters built from the same nested rainbow bands. Crisp printed edges, symmetrical balance, generous negative space, no shadows or gradients, 16:9 landscape.
Can we design legal agent verifiers that are up to 1,000x cheaper? Verifiers are LLM judges that check an agent’s work against rubric criteria: they’re used both in agent benchmarking and as reward signal in post-training. But verifiers can be a bottleneck at scale. For
https://x.com/harvey/status/2061866491033899371
Law professors wrote questions they were asked during office hours. Gemini 2.5 & humans answered them then other law professors blindly judged the results: -Gemini had a 75% win rate vs. professors -Gemini’s answers were rated LESS harmful than humans -Newer models do even better
https://x.com/emollick/status/2061876620638486584
Verifiers are important for scaling evals/RL But costs add up! So can we make them cheaper? Some great work by @Vtrivedy10 @jakebroekhuizen in conjunction with @nikogrupen @gabepereyra and the Harvey team on this
https://x.com/hwchase17/status/2061867746141356427





Leave a Reply