A five-year study of file-system metadata

William J. Bolosky, John R. Douceur, Jacob R. Lorch, Bill Bolosky, John (JD) Douceur, Jay Lorch

ACM Transactions on Storage |

For five years, we collected annual snapshots of file-system metadata from over 60,000 Windows PC file systems in a large corporation. In this article, we use these snapshots to study temporal changes in file size, file age, file-type frequency, directory size, namespace structure, file-system population, storage capacity and consumption, and degree of file modification. We present a generative model that explains the namespace structure and the distribution of directory sizes. We find significant temporal trends relating to the popularity of certain file types, the origin of file content, the way the namespace is used, and the degree of variation among file systems, as well as more pedestrian changes in size and capacities.We give examples of consequent lessons for designers of file systems and related software.