I tried running with a smaller local model. Normal chat was okay, but multi-step ReAct/tool calling became fragile pretty quickly. A lightweight mode that uses fewer tool rounds and falls back to plain explanation when parsing fails would make local deployment much more practical.
I tried running with a smaller local model. Normal chat was okay, but multi-step ReAct/tool calling became fragile pretty quickly. A lightweight mode that uses fewer tool rounds and falls back to plain explanation when parsing fails would make local deployment much more practical.