Follow-up to my earlier posts on getting this MoE running on Battlemage. It started as "can I fully tax four B70s with one big model," and after a lot of testing it turned into four serving configs I'm happy with — a single-stream latency champion, a 2-card option, a high-concurrency config, and a full-precision one — all shipping in a single Docker image where you pick the config at launch. Benchmarked properly (throughput, latency, and capability) and packaged so you can docker pull and serve (or re-run every benchmark) with one Python script. Origin story, the configs, numbers, the interesting engineering, and repro below.
This is a companion discussion topic for the original entry at https://www.reddit.com/r/LocalLLM/comments/1v29ixv/qwen3635ba3b_on_4_intel_arc_pro_b70_vllmxpu_200/