Up front: I build QuantaMind, an open-source local benchmarking tool (Apache 2.0, runs offline, no telemetry). The data below came out of it. Link at the bottom the numbers are the point of the post.
This is a companion discussion topic for the original entry at https://www.reddit.com/r/LocalLLM/comments/1v8bped/toolcall_accuracy_fell_off_at_9k_tokens_on_a/