>_ ANALYSIS
When the cheap model is right for the wrong reason
A home-automation paper has a narrower and more useful lesson than the usual “small models can replace big ones” pitch: the savings come from separating call sites, not from downgrading an entire agent. That distinction matters because the evidence is strongest where the system is structurally uneven — and weakest wher
A home-automation paper has a narrower and more useful lesson than the usual “small models can replace big ones” pitch: the savings come from separating call sites, not from downgrading an entire agent. That distinction matters because the evidence is strongest where the system is structurally uneven — and weakest where a single average would blur that unevenness away.
Kasnesis and coauthors evaluated Wactorz, a deployed open-source home-automation framework, across five LLM call sites on two real installations. Their abstract reports that capability is not ordered the same way at every site, that the best local model is statistically indistinguishable from hosted models at four of five sites, and that routing each site to its best local model reaches 91.8% versus 95.4% at no per-call cost. In live use, keeping only the two generative sites on the hosted model matched all-hosted performance on the user-judged metric, 39/43 against 39/43, while cutting hosted spend to 28%.
The mechanism is straightforward. Wactorz does not ask one model to do one homogeneous task; it makes five different kinds of calls. The first three are short, structured judgments. The last two — planning and code generation — are generative. That asymmetry changes procurement logic. A model that is merely adequate for a label or a device lookup may be a poor bargain for code synthesis, and the reverse can also be true. The practical unit is therefore the call site, not the whole agent.
The paper’s more interesting signal is not the headline average. It is the actuation-specific failure pattern. The abstract says one small model actuated on 87.2% of requests for devices the site did not own, while another refused every request it received. Those are opposite errors, but they point to the same operational problem: aggregate accuracy can conceal whether a model is usable to drive hardware at all. In a home-automation setting, the question is not only whether a model is often right; it is whether it is dangerously eager when it should refuse, or timid when it should act.
That said, the evidence does not justify turning this into a general rule about all agentic systems. The paper itself limits the claim by design: it studies one deployed framework, two installations with different inventories, and prompts that remain specific to Wactorz’s architecture. It also reports non-model failures in the live run, including a framework bug and an LLM client reachability problem. Those failures matter because they show that cheaper routing does not eliminate infrastructure risk and may not even touch the most stubborn sources of error.
A plausible alternative explanation is that the result is mostly a property of this framework and these homes. That is not a weak objection; it is the correct one. The paper’s own live test partly answers it by showing that the offline benchmark predicted the deployed actuation gap and that the gap appeared as exactly one missed case in twenty-six when actuation moved from the frontier model to a local one. But that is still a deployment-specific pattern, not a universal law.
So the actionable implication is narrower and stronger at the same time: if an agent has structurally different call sites, the first savings should come from localizing the short structured ones and reserving frontier hosting for the generative sites that actually need it. That is an allocation decision, not a faith statement about model size. The signal that would strengthen this thesis is another live deployment showing the same shape — little or no loss when structured sites move local, but measurable loss when generative sites do. The signal that would weaken it is the reverse: structured sites becoming the main source of failure once localized.
Source: https://arxiv.org/abs/2610.09021
