Hello, I’ve been having struggle to follow the latest mid sized models. My company recently freed up usage on 2 H200 gpus. I’m wondering which model I can put on them for agentic coding. Context size 256k. And with around 4-10 concurrent users with vllm. But the most often is 4. Very rarely does it go above that.
This is a companion discussion topic for the original entry at https://www.reddit.com/r/LocalLLaMA/comments/1v8c304/how_to_properly_use_2xh200/