I've designed and used quite a few eval harnesses by now. Here's what I keep coming back to: stronger models need less scaffolding. Pile on more constraints and you just throttle them. A simple circuit breaker does the job. Weaker models need the opposite — more harness to keep them on track. But that raises the real cost question: you pick a cheap model to save money, it fails more, and then you wrap it in complex harness logic anyway. Per-task cost ends up not far from just using a strong model.