{"id":1186747,"date":"2026-09-18T15:53:28","date_gmt":"2026-09-18T22:53:28","guid":{"rendered":"https:\/\/www.microsoft.com\/en-us\/research\/?post_type=msr-project&#038;p=1186747"},"modified":"2026-09-18T16:21:32","modified_gmt":"2026-09-18T23:21:32","slug":"capillary-kernels","status":"publish","type":"msr-project","link":"https:\/\/www.microsoft.com\/en-us\/research\/project\/capillary-kernels\/","title":{"rendered":"Capillary Kernels"},"content":{"rendered":"<section class=\"mb-3 moray-highlight\">\n\t<div class=\"card-img-overlay mx-lg-0\">\n\t\t<div class=\"card-background  has-background- card-background--full-bleed\">\n\t\t\t<img loading=\"lazy\" decoding=\"async\" width=\"1401\" height=\"788\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website.jpg\" class=\"attachment-full size-full\" alt=\"background pattern abstract\" style=\"\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website.jpg 1401w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-300x169.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-1024x576.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-768x432.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-1066x600.jpg 1066w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-655x368.jpg 655w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-240x135.jpg 240w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-640x360.jpg 640w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-960x540.jpg 960w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2025\/10\/website-1280x720.jpg 1280w\" sizes=\"auto, (max-width: 1401px) 100vw, 1401px\" \/>\t\t<\/div>\n\t\t<!-- Foreground -->\n\t\t<div class=\"card-foreground d-flex mt-md-n5 my-lg-5 px-g px-lg-0\">\n\t\t\t<!-- Container -->\n\t\t\t<div class=\"container d-flex mt-md-n5 my-lg-5 \">\n\t\t\t\t<!-- Card wrapper -->\n\t\t\t\t<div class=\"w-100 w-lg-col-5\">\n\t\t\t\t\t<!-- Card -->\n\t\t\t\t\t<div class=\"card material-md-card py-5 px-md-5\">\n\t\t\t\t\t\t<div class=\"card-body \">\n\t\t\t\t\t\t\t\n\t\t\t\t\t\t\t\n\n<h1 id=\"exploiting-distributed-shared-memory-in-gpus-for-next-gen-ai-workloads\" class=\"wp-block-heading\">Exploiting Distributed Shared Memory in GPUs for Next-Gen AI workloads<\/h1>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n\n\t\t\t\t\t\t<\/div>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t<\/div>\n\t\t<\/div>\n\t<\/div>\n<\/section>\n\n\n\n\n\n<p class=\"wp-block-paragraph\">Tree-structured prefix sharing is now standard in large language model (LLM) serving: a radix-tree KV cache lets multiple requests reuse common context prefixes and reduces memory capacity pressure. Yet the attention kernels used with these caches usually retain a conventional execution pattern in which each thread block independently loads its KV tiles from global memory. As a result, shared prefixes are still fetched repeatedly, and cross-request reuse is lost at the point where long-context decoding spends most of its bandwidth. Recent prefix-aware kernels reduce this redundancy, but their cooperation is limited by the shared-memory capacity of a single Streaming Multiprocessor (SM), forcing intermediate results that cross SM boundaries through global memory.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><br>We present Capillary, an attention-kernel design that exploits Distributed Shared Memory (DSM) on NVIDIA Hopper GPUs to break this single-SM barrier. Capillary co-locates thread blocks that share KV prefixes within a hardware thread-block cluster and transfers partial attention results directly between SMs, avoiding global-memory round trips. A lightweight DSM-aware cost model selects query grouping and cluster configuration at each decoding step, balancing saved HBM traffic against DSM bandwidth and cluster scheduling constraints<\/p>\n\n\n","protected":false},"excerpt":{"rendered":"<p>Tree-structured prefix sharing is now standard in large language model (LLM) serving: a radix-tree KV cache lets multiple requests reuse common context prefixes and reduces memory capacity pressure. Yet the attention kernels used with these caches usually retain a conventional execution pattern in which each thread block independently loads its KV tiles from global memory. [&hellip;]<\/p>\n","protected":false},"featured_media":1161918,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","footnotes":""},"research-area":[13556],"msr-locale":[268875],"msr-impact-theme":[],"msr-pillar":[],"class_list":["post-1186747","msr-project","type-msr-project","status-publish","has-post-thumbnail","hentry","msr-research-area-artificial-intelligence","msr-locale-en_us","msr-archive-status-active"],"msr_project_start":"","related-publications":[],"related-downloads":[],"related-videos":[],"related-groups":[793670,1145968],"related-events":[],"related-opportunities":[],"related-posts":[],"related-articles":[],"tab-content":[],"related-researchers":[{"type":"user_nicename","display_name":"Srikant Bharadwaj","user_id":41644,"people_section":"Related people","alias":"srbharadwaj"},{"type":"user_nicename","display_name":"Victor Ruehle","user_id":41027,"people_section":"Related people","alias":"virueh"}],"msr_research_lab":[],"msr_impact_theme":[],"_links":{"self":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1186747","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-project"}],"about":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-project"}],"version-history":[{"count":1,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1186747\/revisions"}],"predecessor-version":[{"id":1186792,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-project\/1186747\/revisions\/1186792"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media\/1161918"}],"wp:attachment":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1186747"}],"wp:term":[{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1186747"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1186747"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=1186747"},{"taxonomy":"msr-pillar","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-pillar?post=1186747"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}