# NeuralMesh Data Catalog: Answers at AI Scale

**Author:** Phil Curran

**Published:** September 29, 2026

![Vast black and white image of endless rows of filing cabinets, with one open drawer revealing files.](https://cdn.sanity.io/images/ult5g8gw/production/61ae4c79a22947d67da4256d436625b1dec79e0c-1584x672.jpg)

## TL;DR

Explore how NeuralMesh Data Catalog makes enormous AI namespaces easier to understand, investigate, and manage through queryable metadata.

- Track capacity, ownership, file age, and growth trends.
- Discover datasets and automate workflows through queries, exports, and APIs.
- Compare snapshots to attribute changes and plan capacity.

AI has rebuilt the infrastructure stack from the silicon up, with accelerated compute, 800-gigabit fabrics, and petabytes of storage holding billions of files and objects. All of that infrastructure keeps generating more data: training datasets, checkpoints, model artifacts, inference logs, synthetic data, and the copies pipelines leave behind. Before long, the namespace is enormous, and some very basic questions get surprisingly hard to answer.

Where is my capacity going? What hasn't been read in 90 days? What changed last week? Who owns all of this data? Infrastructure teams ask these questions every day, and there's nothing exotic about them. Traditional filesystems just weren't built to answer them at this scale.

### **Why Simple Questions Get Hard at AI Scale**

A filesystem like Lustre gives you two basic operations: list a directory and describe a file. There's no index that lets you ask a question about millions or billions of files at once. So if you want to know what's there, you have to go look. A script walks the directory tree, examines files one by one, and adds up the results. At a few million files, that may be manageable. At billions, the same approach becomes increasingly expensive.

As a result, these scans don't happen very often. They may run quarterly, cover only part of the namespace, or wait for a maintenance window. The results may then be copied into another system with its own refresh cycle, permissions, and operational overhead. And the moment that scan finishes, the namespace starts changing again.

The real problem isn't that a query takes too long. It's what happens when infrastructure teams can't easily answer basic questions about the data they're responsible for. Capacity gets added because it's easier than figuring out what can be reclaimed. Data that might be cold stays on expensive storage because nobody can confidently identify it. Engineers open tickets to locate datasets. Unexpected growth is difficult to attribute to a team or workload. Capacity planning becomes part analysis, part guesswork.

AI is still changing the shape of the problem even further. As AI applications become more agentic, they can initiate more work, interact with more data, and generate more activity without a person explicitly starting every task. That makes already-large AI environments more dynamic and harder for infrastructure teams to understand.

Knowing what's in your storage, where capacity is going, and what changed becomes more important as the systems using that data become more autonomous. At AI scale, storage needs to make those answers easier to find.

### **WEKA NeuralMesh Already Knows**

WEKA® [NeuralMesh™](/product/neuralmesh) doesn't have to repeatedly crawl the namespace to find out what it contains. It already tracks the files and metadata it manages. NeuralMesh Data Catalog makes that information available to query. And because it's the same namespace underneath, Data Catalog indexes metadata whether data was written as files or as S3 objects. There's no separate crawl per protocol.

Data Catalog builds its index from a NeuralMesh snapshot. The first pass establishes a baseline. After that, a diff list identifies exactly what changed between consecutive snapshots, and only those changes are added to the index. There's no need to walk the entire namespace again. You choose the interval as well as how much history to retain. The index lives in its own NeuralMesh filesystem and contains metadata only, never file contents.

That gives infrastructure teams a current, queryable view of their data without introducing another system that has to continually rediscover what's already inside the filesystem. And it makes some of those basic questions much easier to answer.

### **Where Is My Capacity Going?**

Capacity Usage Reports show where storage is being consumed across the namespace. Start with a sunburst view of your largest directories and drill down to the leaf. A question like "where did 400 terabytes go?" becomes something you can investigate in a few clicks.

You can also look at capacity by file extension, file size, and owner. That makes it easier to see when, for example, one team's intermediate output has grown larger than the training dataset that produced it.

File age adds another useful dimension. You can see how much capacity hasn't been accessed recently, turning "we probably have a lot of cold data" into something measurable.

Trend analysis puts that information on a timeline and projects it forward, giving infrastructure teams a clearer picture of how quickly capacity is growing and where that growth is coming from.


![NeuralMesh Data Catalog File and Object Use Chart](https://cdn.sanity.io/images/ult5g8gw/production/fecb90a1bd71ab95c6ddf5d33be87f56d4b9c084-4800x3000.png)


### **What Data Do I Have?**

At AI scale, finding a specific dataset can feel like looking for a needle in a planet of haystacks. Sometimes you know what you're looking for, but not where it lives. Discovery lets you query the Data Catalog index using attributes such as path, size, owner, modification time, and access time.


![NeuralMesh Data Catalog Metadata Query](https://cdn.sanity.io/images/ult5g8gw/production/7636bae1c199a10a79db44519a70bfbe17d251d1-4800x3000.png)


You can start with a template looking for things like files not accessed in 90 days, the largest files in a directory, everything matching a particular extension or build your own query.

You can also start from the visualizations themselves. Click a segment of a chart and Data Catalog opens the query behind it, already populated and ready to refine. Instead of treating a chart as the end of a report, you can use it as the beginning of an investigation.

Results can be exported to CSV or accessed through the REST API. The UI is built on that same API, so the information an administrator can explore interactively can also be used programmatically. That's important because understanding the data is only the first step. Once attributes like location, size, age, ownership, and access patterns are queryable, they can also become inputs to workflows and automation.


![NeuralMesh Data Catalog API Query](https://cdn.sanity.io/images/ult5g8gw/production/46fdcbab1abd64785d95744956336fdc71c91670-4800x3000.png)


### **What Changed?**

By comparing any two snapshots, Data Catalog shows you what changed and when. View which files were added, modified, or deleted, along with the capacity impact of those changes. It also ranks directories by storage impact. So if capacity jumps unexpectedly over a weekend, you can trace that growth back to the directories responsible. Maybe a training run left behind a large set of checkpoints. Maybe someone created another copy of a dataset. Maybe a new workload is growing much faster than expected.

Over time, that same change history shows how much each directory actually changes, which is useful when you're sizing replication or planning capacity for the next quarter.

Instead of discovering the increase at the end of the quarter, you can see what changed and where it happened. That changes the capacity conversation. Growth becomes something infrastructure teams can investigate and attribute, rather than simply accommodate.


![NeuralMesh Data Catalog Comparison Insights](https://cdn.sanity.io/images/ult5g8gw/production/d33b5c731f675ac9d3d3d9315f7f20aaa1d6876c-4800x3000.png)


### **Storage Should Know What's in It**

We built NeuralMesh because AI required a different storage architecture. As AI infrastructure evolves, what we expect from storage has to evolve too. Today, Data Catalog helps infrastructure teams answer practical questions about enormous namespaces: What data do I have? Where is my capacity going? What's cold? What changed?

But making that information structured and queryable opens up a bigger possibility. Imagine infrastructure that can use attributes about the data itself to help determine what happens next. Which data needs to be on high-performance storage for an upcoming workload? Which data hasn't been touched and can move somewhere else? Which new files should trigger another step in a workflow?

That's where a catalog can become more than a way to understand your data. The information it provides can become an input to how that data is managed. We're at the beginning of that journey. NeuralMesh Data Catalog starts with something fundamental: making the data inside enormous AI environments easier to understand. Because before you can automate what happens to your data, you have to know what's in there.

NeuralMesh Data Catalog is available today as part of your NeuralMesh subscription. Enable it on the filesystems you choose. No new license. No new vendor. And no separate system holding a copy of your metadata.
