プロジェクト
Large language model (LLM) inference has become an operational workload that fleets must provision, monitor, and bill against latency service-level objectives (SLOs). A single throughput number cannot answer the questions operators actually face, because generative serving couples two phases with…
Tree-structured prefix sharing is now standard in large language model (LLM) serving: a radix-tree KV cache lets multiple requests reuse common context prefixes and reduces memory capacity pressure. Yet the attention kernels used with these caches usually retain a conventional…
Agents for Productivity (A4P) is a M365 Research initiative to enable Microsoft to deliver reliable, highly capable, and scalable agentic solutions that drive measurable productivity impact. The strategy addresses two core challenges: technological gaps (tool integration/selection, memory & context management,…