Tool-call accuracy fell off at ~9k tokens on a model whose context window is 16k and memory could have held 53k

Up front: I build QuantaMind, an open-source local benchmarking tool (Apache 2.0, runs offline, no telemetry). The data below came out of it. Link at the bottom the numbers are the point of the post.


This is a companion discussion topic for the original entry at https://www.reddit.com/r/LocalLLM/comments/1v8bped/toolcall_accuracy_fell_off_at_9k_tokens_on_a/