The Data Catalog Features That Boost Productivity and Organization

Emma Vandermey
Business meeting and presentation in modern conference room for colleagues

Table Of Contents

The greatest obstacle toward becoming a data-driven company is organization. The more data you have, the harder it is to organize. Without basic organization, data consumers are stuck with the messy and arduous chore of wading through massive piles of data. Worse, if they’re not data engineers or data professionals, they may feel so overwhelmed they  just give up.

Enter the data catalog.

What is a data catalog?

A data catalog is an organization system for massive piles of data and data assets. Instead of wading through endless tables, cryptic column names, potentially-sensitive data, and schemas with no guarantee of structure, a data consumer can use a data catalog to discover, search, and use the data they need.

Data catalogs allow for efficient data management across the entire data ecosystem. Features and capabilities may vary depending on whether it’s an open source data catalog or paid data catalog software, but you can count on data catalogs to be compatible with multiple data sources, data warehouses, and data platforms. They’ll also handle metadata management, work within a data governance framework, and solve for a variety of use cases at both enterprises and SMBs.

How a data catalog works?

A data catalog is a centralized repository that provides metadata about an organization’s data. It doesn’t store or manage the data itself, but a catalog can tell you all about your data. People often relate data catalog solutions to:

  • Librarians: They know where to find all the books, but they don’t know what’s in them
  • Maps: They show you where all the data lives across your landscape
  • The Rosetta Stone: It translates the different languages of data into a universal language
  • Treasure maps: They show the hidden gems that can improve business decisions

In other words, data catalogs point you to data and tell you all about it. They just don’t know what’s actually in the data. For this reason, they’re often called “metadata repositories” or even “data dictionaries.”

How metadata is used in a data catalog

Metadata describes the data across an entire data ecosystem and gets stored and managed in a data catalog. Metadata includes:

  • Data sources: SaaS apps, internal product databases, and cloud storage services (e.g. AWS S3)
  • Data destinations: Data warehouses, data lakes, and BI platforms
  • Data lineage: The history of how data is created, processed, and used
  • Data quality: The degree to which data is accurate, complete, and consistent
  • Data usage: The way that data is being used by consumers and acquirers
  • Data policies: The rules that govern how data is used

Metadata can be used not just to describe data, but also to relate datasets. For example, a data catalog can tell you that one dataset is a history of consumer transactions that can be joined with a dataset of demographic information. Related datasets are often called “collections,” which users or the data catalog can create.

When the catalog contains a sufficient amount of metadata, that metadata is indexed, or optimized for search and discovery. Data assets can also be indexed in a process called “data asset indexing,” which is the process of organizing and categorizing data assets within a catalog or repository. It involves creating an index or a searchable catalog of available data assets so it’s easier for users to locate and access specific datasets based on their search criteria.

Exploring, searching, and discovering metadata in a data catalog

One of the primary data catalog use cases is searching for data sets. This is where indexing comes into play. Indexes can be built based on various aspects of the metadata, such as the types of data, keywords, tags, or other metadata-based criteria.

Depending on how the data is indexed, consumers can search and discover data sets that pertain to their specific use cases. For example, a data team might produce an API for upstream supply chain vendor inventories and tag the API as a data source for supply chains. Later, a finance team might search the catalog for the keyword “supply chain” and discover the API.

Catalogs don’t take usability into account, meaning they have no sense of whether the finance team can actually use an API. A catalog allows a user to search a metadata repository and responds back to the user with a “yes” or “no” as to whether there’s data available that’s relevant to their search criteria.

Metadata standards can be enforced with automation and governance to increase the ability for users to search and discover as well. For example, when searching for a pair of men’s shoes on Amazon, you can refine your search with metadata related to:

  • Manufacturer brand (e.g. Adidas, Nike)
  • Shoe size (e.g. sizes 4-19)
  • Shoe width (e.g. XX-Narrow through XX-wide)
  • Running shoe support type
  • Outer materials (e.g. canvas, leather, vinyl)
  • Color

Similarly, a data catalog can present refinable metadata classifications for business data based on how the business has established a metadata framework. Premium enterprise data catalogs like Dremio can present a unified view of all data assets across multiple systems and locations (a feature known as “data virtualization”). In other words, one user interface for locating data, even if it’s in 35 data centers across the globe.

Data catalog features

Many data catalog products feature a unique variety of capabilities. Here are some of the most important features you should consider when looking at data catalogs:

  • Automated data discovery
  • Data lineage
  • Data collaboration and sharing
  • Data monitoring and anomaly detection
  • Metadata curation
  • Data customization

Automated data discovery

Most enterprise data ecosystems are—for a lack of a better word—huge. They typically have dozens, if not thousands, of data sources. It’s too much work to manually configure a data source to be cataloged. This is where automated data discovery can make all the difference.

Though any organization can take advantage of automated data discovery, it’s the larger companies that benefit most. When they want to improve their data governance and make data more discoverable, usable, and compliant, the data catalog can automatically scan all known data sources and extract what metadata it can detect. It can also analyze raw data assets to identify their characteristics and make intelligent decisions about classifying the data at the source. This includes making assumptions about the underlying data structures and any patterns it might contain.

By removing the manual labor from this process, there’s less room for error and much more data can be analyzed in a short period of time (as opposed to a human reviewing and classifying the data). In the process, some catalogs can also look through and understand the lineage of data assets, checking for errors, completeness, and consistency.

Obviously, automating this work is a significant time saver, especially in ecosystems with hundreds or thousands of data sources. The faster data can be cataloged, the faster consumers can use it.

Data lineage

All data has an origination time, place, and reason. Data lineage gives you the who, what, where, when, and why behind the existence of every piece of data across the ecosystem. In addition to tracking creation, lineage also allows you to see how data has changed over time.

For example, let’s say your organization has a CRM that stores customer data. Let’s also say your organization’s data privacy team wants to make sure the data is used in a compliant and secure manner. A catalog’s lineage tracking can show how the data was created, processed over time, and even used—whether by a person, system, or a process.

Your organization can use this information to:

  • Identify potential compliance risks
  • Improve the accuracy and completeness of the data
  • Investigate data quality issues
  • Optimize data usage

This is especially important for regulated industries dealing with PII (e.g. banking, retail) and PHI (e.g. healthcare, insurance). Many compliance programs require audit logs of data access and usage. The data lineage capability of a catalog can fulfill this requirement.

If the benefits of data lineage tracking sound expensive, rest assured that even open source data cataloging tools have this feature now.

Data collaboration and sharing

Data catalogs aren’t merely a portal into metadata; they also allow for collaboration. Data collaboration within a catalog involves knowledge sharing among users who interact with the catalog. It provides a platform for individuals and teams to collaborate, discuss, annotate, and provide feedback on data assets. Here are some key aspects of data collaboration in a data catalog:

  • Comments and discussions: Users can engage in discussions, leave comments, or ask questions about specific data assets within the catalog 
  • Annotations and tags: Users can annotate data assets with additional information or insights Annotations can include descriptive notes, data quality assessments, or usage recommendations. 
  • Ratings and reviews: Users can rate and provide feedback on the quality, usability, or relevance of specific data assets. These ratings and reviews help others in the organization to assess the value and suitability of the data for their needs
  • Data requests and sharing: The data catalog can facilitate requesting access to specific data assets. Users can submit requests to data owners or administrators, specifying the purpose and duration of access 
  • Collaboration spaces and workflows: Advanced data catalogs may offer collaboration spaces or workspaces where teams can collaborate on specific projects or datasets
  • Data lineage and documentation: Collaboration in a data catalog can extend to data lineage and documentation. Users can contribute to documenting the lineage of datasets, capturing information about data transformations, data sources, and data dependencies 
  • Notification and subscription: Users can subscribe to specific data assets or datasets of interest, enabling them to receive notifications or updates whenever there are changes or new information related to those assets. This ensures that users stay informed about updates and can actively participate in data discussions or collaborations 

For organizations that want to become more data mature and capable, collaboration is essential. The more people can contribute, the more trustworthy the data becomes. When data is trustworthy, people can confidently solve business problems.

Data monitoring and anomaly detection

Data monitoring and anomaly detection can help organizations ensure the quality and integrity of their data. Data monitoring involves tracking changes to data over time, while anomaly detection identifies unusual or unexpected patterns in data.

In a data catalog, data monitoring and anomaly detection can:

  • Identify and correct errors in data: Tracking changes to data over time can identify errors introduced into data by comparing the data’s current state to its historical state
  • Detect anomalies in data: Anomaly detection can identify unusual or unexpected patterns in data to identify potential problems, such as fraud or data corruption
  • Improve data quality: Identifying and correcting errors in data can improve data quality, making it more reliable and useful for analysis
  • Reduce risk: Detecting anomalies in data reduces overall business risk by identifying potential problems with data before they cause problems for the organization

Metadata curation

Metadata curation is the process of organizing, cleaning, and enriching metadata in a data catalog. It is an important part of data governance and can help to improve data discovery, usability, and compliance. Another name for metadata curation is “metadata management.” Metadata management tools organize, clean, and enrich metadata.

One more approach to implementing metadata curation in a data catalog is to use a manual process. This can be time-consuming but effective in ensuring that metadata is accurate and complete.

The best way to approach this is with automation, potentially involving AI/ML. Some catalogs have this capability built in. 

The main benefit to metadata curation is to ensure that your catalog presents no risk to your overall governance and compliance policies. In the process, it will also make data more trustworthy and usable across your organization.

Data customization

Data customization is the process of tailoring the catalog to the specific needs of a particular user or group of users. This can be done by adding or removing metadata, creating custom views, or generating reports.

As is the case with any technical business tool, there will be light users and heavy users. Those who use the catalog most will want customization features that save them time and frustration. Catalogs are great at presenting data assets, but not typically great at making them usable.

Over time, as the catalog matures with customization, more people will want to use it, which will help drive a company towards data maturity and data-driven decision making.

Revelate and data cataloging tools

Companies often purchase data catalogs with the hope that the catalog will increase data sharing across lines of business and functions. Unfortunately, that’s almost never the case.

Imagine going to the library and asking a librarian for a book. The librarian says, “Yes we have that book, but it’s only in Italian.” If you don’t read or speak Italian, you will probably have trouble getting value out of the book. The librarian did their job and told you that the book is in stock, but can’t do anything to make it readable for you.

Revelate closes that gap.

Whereas a catalog says, “Yes, the data you want exists and we can tell you where to get it,” Revelate actually makes data acquisition an entirely self-service process. You can search for what you need and get access to it right away in a format that you can actually use.

Catalogs focus on organization and indexing. Revelate leverages data catalogs to provide data products that you can use right out of the proverbial box. We have tight integrations with many data catalogs so you can go from “organized” to “productive” in a short time.

Revelate provides data products with metadata designed specifically to help people find what they need. We provide everything you need to transform raw data assets into consumable data products.

Revelate makes catalogs twice as useful

Data catalogs are excellent tools for organizing your data, tracking its history, ensuring a baseline level of quality, and providing trusted data to the people who need it. However, most data consumers don’t want data assets—they want data products.

As organizations navigate the complexities of their data ecosystems, the importance of data catalogs organizing data cannot be overstated. Yet, cataloging and indexing data is no longer enough to unlock the full potential of data. This is where Revelate steps in, revolutionizing the way organizations harness the power of their data assets.

Revelate goes beyond the traditional role of data catalogs by transforming raw data assets into actionable data products. By leveraging the rich metadata foundation established by data catalogs, Revelate empowers data consumers with self-service access to data that is immediately usable and tailored to their specific needs.

While data catalogs lay the foundation for effective data organization, Revelate’s transformative approach takes data utilization to new heights, revolutionizing the way organizations extract actionable insights from their data assets. By seamlessly integrating with data catalogs and providing customizable data products, Revelate empowers data-driven organizations to thrive in an increasingly complex and data-rich environment.

Unlock Your Data's Potential with Revelate

Revelate provides a suite of capabilities for data sharing and data commercialization for our customers to fully realize the value of their data. Harness the power of your data today!

Get Started