
The partnership addresses a persistent gap in enterprise AI adoption: the majority of any given workforce cannot use AI tools because the data those tools would need to work with cannot legally or safely be sent to a hosted third-party inference endpoint. Kasm AI Workspaces on Intel Xeon 6 with AMX changes that calculus by running local models — including the distributed and mixture-of-experts (MoE) LLMs that Intel and its ecosystem are advancing — inside each user’s ephemeral workspace container. Recent open-weight models running on this architecture, including Qwen3-Coder-30B-A3B, now approach the capability of leading frontier models on the workloads that dominate day-to-day enterprise use: chat, retrieval-augmented generation, tool calls, and code assistance.
“This partnership with Intel changes what’s possible for enterprise AI. By running Kasm AI Workspaces on Intel Xeon 6 with AMX and OpenVINO, we give organizations a way to deploy real AI to every desk — chat, code assistance, retrieval-augmented generation, and autonomous agents — without any data crossing the perimeter or touching a third-party inference endpoint. That is what workforce-scale AI adoption actually looks like, and it is what regulated industries have been waiting for.”
Jaymes Davis, Chief Technical Evangelist at Kasm Technologies
Beyond CPU-based inference, the joint architecture provides a single workspace control plane that orchestrates AI workspaces across Intel’s full compute portfolio — CPUs with AMX, integrated NPUs, and discrete GPUs — and delivers each session to any browser on any endpoint. For workloads that demand GPU acceleration, Kasm 1.19 supports SR-IOV bifurcation of Intel Arc Pro cards, allowing a single physical GPU to serve multiple isolated workspaces as virtual functions. The result is a unified deployment path for private AI applications that spans everything from lightweight interactive chat to long-form document generation and autonomous coding agents. Regulated industries — healthcare, finance, legal, defense, and government — have been early adopters, with the architecture reaching cost parity with per-seat AI subscriptions at approximately 40 provisioned users per node and inverting favorably above that threshold.






