The Data Systems group at Microsoft Research works on problems in data management. Our current areas of focus include infrastructure for large-scale cloud database systems, storage and indexing in key-value stores and vector databases, leveraging modern hardware to accelerate performance in database systems, query optimization, reducing the total cost of ownership of databases through auto-tuning, and enabling flexible ways to discover, transform and clean data sets containing both structured and unstructured data.
Our research has had a significant impact on the industry. Technology developed in our projects has shipped in several Microsoft products and services and found wide-spread adoption when released to open source. Examples of these technologies are: physical design tuning in the Database Tuning Advisor (opens in new tab) in Microsoft SQL Server, flexible resource allocation techniques in Azure SQL Database cloud infrastructure, fuzzy matching and fuzzy deduplication in Power Query (opens in new tab), Azure Data Factory (opens in new tab), Dynamics 365 (opens in new tab), Microsoft SQL Server Integration Services (SSIS) (opens in new tab), data transformation by-example in Power Query (opens in new tab), lock-free indexing for Microsoft SQL Server’s in-memory OLTP engine (“Hekaton”) (opens in new tab), technology for enabling real-time analytics in Microsoft SQL Server (“Apollo”) (opens in new tab), fast parsing of CSV and JSON data in Azure Synapse (“Mison”), the Trill event processing engine in Azure Stream Analytics (opens in new tab), the FASTER key-value store which is open source on GitHub (opens in new tab) and is used by Azure Stream Analytics and Azure Durable Functions, the Garnet cache-store which is open-source on GitHub (opens in new tab) and is deployed by Azure Resource Manager, Azure Resource Graph, and as a managed service (preview) (opens in new tab) from Azure Data, the Orleans open source actor framework (opens in new tab) for building distributed applications, and the mapping compiler for Microsoft’s open source Entity Framework (opens in new tab).
Our research has also had significant impact on the academic community. We have published in top conferences such as ACM SIGMOD, VLDB, IEEE ICDE, CIDR, ACM SIGKDD, ACM SIGIR, The Web Conference, NeurIPS, etc. Our work has resulted in two VLDB 10-Year Best Paper Awards, an ICDE Influential Paper Award as well as Best Paper Awards at ACM SIGMOD, VLDB, IEEE ICDE and CIDR.
NEW: Awards for Data Systems Group papers at VLDB 2026 and ACM SIGMOD 2026
- Our work on Garnet: A Next-Generation Cache-Store for Accelerating Applications and Services received the Best Paper Award at the 52nd International Conference on Very Large Data Bases (VLDB 2026) (opens in new tab). More details on Garnet can be found here.
- Our paper Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking and Optimization (opens in new tab) received the Best Paper Honorable Mention at VLDB 2026.
- For an overview of Microsoft’s work presented at VLDB 2026, see Microsoft at VLDB 2026 (opens in new tab).
- The paper CoddSpeed: Hardware Accelerated Query Processing in Microsoft Fabric (opens in new tab) received the Best Industry Paper Award at ACM SIGMOD 2026 (opens in new tab). This project is a collaboration across several teams in Azure Data as well as our group.