Skip to main content
Data Management for Machine Learning

Machine Learning Experimentation and Metadata

Published: 2026-08-07
Level: postgraduate
Audience: Postgraduate students in Machine Learning and Data Management

Prerequisite Knowledge

This lecture builds on the following concepts from earlier lectures. If any feel unfamiliar, review the linked notes before proceeding.

Previously Covered in This Subject

  • Data integration and ingestion — covered in Lecture 9 (Data Integration and Data Transformation) and Lecture 5 (The Modern Data Stack and Data Pipelines)
  • Metadata fundamentals — covered in Lecture 2 (Query Paradigms, Storage Architectures, and Data Pipelines)
  • Data pipelines and the modern data stack — covered in Lecture 5 and Lecture 10 (Orchestration, Automation, and Version Control for Data Pipelines)
  • Train, validate, and test partitioning — covered in Lecture 7 (The Machine Learning Lifecycle)
  • Cross-validation and hyperparameters — covered in Lecture 7 (Model Development, Algorithm Selection, and Hyperparameters)
  • Model registry and lineage — covered in Lecture 7 (Feature Stores, Model Registry, and the Drift Feedback Loop)
  • Data lineage — covered in Lecture 8 (Data Ingestion, Validation, and Data Quality)
  • Data quality dimensions — covered in Lecture 8 and Lecture 9 (Data Integration and Data Transformation)


11.1 Opening Story — Attachment, Mindset, and Ownership

Every class here begins with a "positive point" — a short story the whole room discusses before the technical content. This class opened with a picture of a man standing in front of his house as it burned, and the room was asked what the picture says. The three readings that came back are worth keeping in mind, because each one is a different mindset at work.

Hook: What do you see when a man watches his house burn? Three answers came from the room. The first reading: a man watching the house he built over time go up in flames in front of his eyes — grief at a life's work lost. The second reading flipped the scene: maybe he is the person who set the fire — "how clever I was to have bought that house" — a destructive mindset. A third student read it as an experiment: he is testing which petrol or which oil burns the house fastest. One picture, three mindsets.

The reveal tied the picture to the day's topic: the class is about ML experimentation and metadata, so in a sense the man is running his own experiments — he is building his own metadata. Every experiment produces observations, and the way you record, compare, and interpret those observations is the metadata of your work. The story that followed made the same point about the mind that runs the experiments.

11.1.1 The Burning House Story

The full story: in a city stood a luxurious house, considered the most beautiful house in the city. People could not pass it without praising it. The owner left the city for a few days on work. When he returned, he saw smoke coming out of his house — flames were rising from his beautiful house. A crowd of spectators gathered.

The instructor compared the crowd to drivers on a highway who slowly stop their vehicles to see what is happening after an accident: it is a curious mind, and very few people actually go and help. The owner pleaded for help, and nobody helped him. He panicked, watching his house burn and thinking what to do now.

Then his eldest son arrived and said: "Papa, don't panic. Everything will be alright." The father asked why he should not panic — his house is burning. The son explained: a few days earlier, while the father was out of the city, he found a great buyer for the house, who proposed three times the value of the house (property in India always appreciates — minimum two times, minimum three times). The deal was so good he could not refuse, so he finalized it without the father's consent. Hearing this, the father worried less, breathed a sigh of relief, and stood there watching the house burn — he had become a spectator like everyone else.

Then the second son came and asked what the father and brother were doing — the house is burning and they just stand there watching. The father explained that the elder brother had sold the house at a very good price, so it is no longer their house. The second son replied: the deal has not yet been confirmed — they have not got the money yet; nobody knows when the buyer will pay, and how will he pay for a burning house? The father grew worried again, started thinking about how to control the fire, and once more pleaded with the people gathered around for help.

Finally the third son arrived. He had just come from the house of the buyer. He told the buyer about the fire, and the buyer said he never goes back on his words: no matter what has happened, he will still buy the house and the land and rebuild it, and he will give the money. The father was relieved again and went back to watching the house burn.

Picture the scene as a chart with a single line: the father's mood. The flame height is roughly constant, but the father's state swings from panic to relief to panic to relief, driven only by news arriving from outside. The fire is the same fire; only the information reaching him changes.

Pitfall: attachment makes your mood a passenger. The father's state flipped four times while the house burned at the same rate: panic at first sight, relief at the news of the sale, panic at the unpaid deal, relief at the buyer's promise. Nothing he did changed the fire. If your happiness depends on news you cannot control, your emotional state swings with every message — the same trap appears in data work when a single failed experiment ruins a day.

11.1.2 The Moral — Thinking, Attachment, and Sorrow

A student read the moral aloud: while the house was burning, its master's thinking changed many times, and because of that his behavior changed. When we get attached to something and it is taken away, we feel sad. When we look at something that is not related to us, we feel a different freedom, and sorrow does not even touch us. So being sad or not depends completely on our thinking and mindset. By controlling our thinking, or giving it the right direction, we can avoid many sorrows and troubles and reach new heights in life.

The instructor extended the point: meditation teaches the same idea — the body is not yours, nothing we brought into this world is truly ours, so do not get attached emotionally. That does not mean becoming a cold spectator of other people's pain: feel happy for the good things happening around you, feel sad for the bad things, help a person who is hurt or pray for their quick recovery and healing. The point is to loosen attachment so that external events do not own your mood.

The instructor also drew on a folk saying — if I get blood, it is my blood; if you get blood, it is tomato ketchup — because we treat our own pain as critical and another person's pain as trivial. The wise mind keeps perspective and responds to both with care.

Recap: sadness follows attachment. Whether a loss touches you depends on how you hold the thing you might lose. Control the direction of your thinking, and the same event produces a very different emotional result — the moral of the house story, and the entry point for the day's lesson on disciplined experimentation.

11.1.3 Ownership at Work

That same morning the instructor had written the opposite message to a work team: nobody was taking responsibility for the product, people only responded when they were tagged, and items sitting in the backlog were not being picked up. The message was about ownership and attachment — "this is a product that we own, see how it is helping all our clients."

The story and the work message are two sides of one idea: attachment to outcomes is what creates suffering, but ownership without emotional attachment is the healthy middle — caring deeply about the work while keeping the mind steady when results change.

Intuition: think of a goalkeeper. A goalkeeper who fears conceding plays timidly and lets in easy goals; a goalkeeper who is careless about conceding stops trying. The strong goalkeeper owns the goal — dives, commits, and stands up again after every miss. Ownership without attachment means full commitment to the attempt, and full calm about the outcome. In machine learning terms this is exactly the mindset of disciplined experimentation: run the model, watch the result, adjust, repeat — without panicking at any single result.

11.1.4 The Cockroach Theory — React vs Respond

The instructor's favorite bonus story, popularized by Sundar Pichai, is the cockroach theory of self-development. At a restaurant, a cockroach suddenly flew from somewhere and sat on a lady. She started screaming out of fear, and with a panic-stricken face and trembling voice she jumped with both hands, desperately trying to get rid of the cockroach. Her reaction was contagious — everyone in her group also turned panicky. She finally managed to push the cockroach away, and it landed on another lady in the group, who continued the drama. The waiter rushed forward to their rescue, and the cockroach landed on the waiter. The waiter stood firm, composed himself, and observed the behavior of the cockroach on his shirt. When he was confident enough, he grabbed it with his fingers and threw it out of the restaurant.

The reflection: was the cockroach responsible for the ladies' histrionic behavior? If so, why was the waiter not disturbed? He handled it almost perfectly, without any chaos. It is not the cockroach, but the inability of those people to handle the disturbance caused by the cockroach, that disturbed the ladies.

The same applies to life: it is not the shouting of a father, a boss, or a wife that disturbs us, but our inability to handle the disturbance caused by their shouting. It is not the traffic jam on the road that disturbs us, but our inability to handle the disturbance the traffic jam creates. More than the problem, it is our reaction to the problem that creates chaos in our lives.

Pitfall: reacting instead of responding. Reactions are always instinctive — the lady jumped before she thought. Responses are always well thought out — the waiter observed, waited, and then acted. Three common traps: treating the trigger (the cockroach) as the cause of the disturbance, letting one person's panic spread through the whole team, and confusing speed with decisiveness. The waiter was fast; the ladies were merely loud.

The lesson: do not react in life — always respond. The women reacted, whereas the waiter responded. A person who is happy is not happy because everything is right in his life; he is happy because his attitude toward everything in his life is right. This maps directly onto the day's topic: machine learning experimentation is the structured version of "respond, don't react" — every experiment is designed, observed, measured, and compared before any decision is made.

Recap + bridge: ownership without attachment keeps the mind steady; respond, don't react, keeps the work clean. The rest of the class applies both lessons to machine learning: experimentation gives you a structured way to respond to model results, and metadata gives you the memory to compare experiments without panic.

Q: What do you see in the picture of a man watching his house burn? A: Several readings came up: a man watching the house he built over time burn in front of his eyes; the man himself as the arsonist, admiring how clever he was to have bought that house (a destructive mindset); or an experimenter testing which petrol or oil burns the house fastest. The reveal: today's topic is ML experimentation and metadata — in a way, he is building his own metadata through experiments.

11.2 Flashcard Recap — Data Integration, Ingestion, and Pipelines

The technical part of the class opened with a flashcard: a rapid-fire recap of the previous classes on data pipelines and data integration. Every answer here builds the foundation for the metadata content of the day — as the instructor put it, the recap "will give a base for our metadata," because transformation rules and quality checks all rely on good metadata. If you can say what each of these terms does in one sentence, you are ready for what comes next.

11.2.1 The Primary Goal of Data Integration

Data integration is the discipline of bringing data from many places into one consistent picture. Without it, each department holds its own copy of reality, and the copies disagree. Integration is what turns scattered sources into a single version of the truth that downstream applications can rely on.

Q: What is the primary goal of data integration within an organization? A: You get data from multiple sources and combine them into a single unified view, making sure quality and consistency are taken care of. From that single source you can create your data mart, build pipelines from there, and create applications on top. The goal is to combine various data types, different formats, and different sources into a unified fact base for analytics, which will be useful for downstream applications and further analytics.

Notice the words used in that answer: unified view, quality, consistency, fact base. Integration is not just copying — it is reconciling.

11.2.2 Ingestion vs Integration and Data Quality

A unified data set drives better business performance because it gives a holistic view of all the possible insights; that boosts the business decisions you want to make and enhances performance. Applications — BI or ML — are built for certain business needs. When the data you are interested in is curated in one place through integration, it becomes easy to query it from a BI platform.

Q: How does a unified data set drive better business performance? A: It gives you a holistic view of all the possibility of insights, which gives a boost to your business decisions and enhances performance. Applications are built for certain business reasons — BI or ML — and when the data is curated somewhere through integration, you can simply query it from a BI platform.

In the context of data movement, the specific role of a data pipeline is to define the path through a technical system, so that the data can be fabricated — the pipeline lets you understand the structure and the meaning of the data as it moves. Fabricated here means "shaped into its final, useful form" — the pipeline does not just carry bytes; it gives them structure and meaning step by step.

Q: What is the specific role of a data pipeline with respect to data movement? A: We define the path through a technical system so that we can fabricate the data — we make sure we can understand the structure and the meaning of the data.

Ingestion is the first leg of that journey — getting the raw material in. The distinction between ingestion and integration matters because the two stages have different goals and different quality expectations.

Q: What does data ingestion mean? A: Reading data from various sources and putting it into a landing zone — collecting and transferring data from different sources into one target storage system. If the pipeline is automated, it also does quality checks automatically during ingestion.

The difference in quality behavior is the key contrast between the two stages: one collects, the other curates.

Q: How does data integration differ from data ingestion on data quality, constraint checks, and the rest? A: Ingestion just collects everything — it does not improve quality unless you automate checks. Integration improves the data quality through transformation: lots of merging, lots of filtering, removal of noise from the data and of irrelevant features.

The side-by-side picture:

Dimension Data ingestion Data integration
Goal Move data from sources into one target storage Combine sources into one unified, consistent view
Quality No improvement unless checks are automated Actively improves quality: merging, filtering, removing noise
Context The context of use is not set yet Knows what the downstream applications need
Expertise needed Lower Higher coding and domain expertise
Example act Copy a file into the landing zone Join, cleanse, and enrich it for a model

Q: Why do data integration pipelines require higher coding and domain expertise than ingestion pipelines? A: You need expertise to keep the features that are relevant for your business model. Suppose you have a mix of financial and healthcare data and you are creating a healthcare model — the financial details are not needed at all. At the ingestion layer the context is not set, but at the integration layer you know what the applications need, so it is your responsibility to get the data from the ingestion layer, curate it, and make it ready for analytics.

That last point is the bridge to metadata: because we are doing this data transformation, we need good metadata; metadata lets us define the rules, and so it gives a good quality check. You cannot enforce a rule you have not recorded.

11.2.3 Pipelines, Real-Time Data, and Agile Practices

Modern business wants decisions in real time, and that demand changes what pipelines must carry. Data no longer arrives in one monthly batch; it is transformed in flight, from many places at once.

Q: What modern challenge is driven by the need for real-time decision making? A: Continuous data — data transformed in flight, in multiple places, lots of time-series data, and a demand for continuous data. That demand for always-on self-service data comes from DevOps and Agile — and ML Ops belongs there too.

Agile project management is built on continuous integration and CI/CD pipelines: flexible plans, high uncertainty, embracing changes, high customer interaction, self-organized project teams, data that comes and goes, continuous integration everywhere. When the delivery cadence speeds up, the data supply chain has to keep up — so data is always on, self-service.

Q: Which practice has increased the organization's demand for always-on self-service data? A: DevOps and Agile — you can add ML Ops as well. In Agile we have continuous integration and a CI/CD pipeline: flexible, high uncertainty, embrace changes, high customer interaction, self-organized project teams, and continuous integration.

11.2.4 What Data Integration Tools Enable

Integration software is judged by what it can do with the data once it is collected. A widely used catalog of these tasks comes from Gartner, the research and advisory firm.

Q: According to Gartner, what are the primary tasks enabled by data integration software products? A: Data transformation, and before that data cleansing; then loading and analytics; and data enrichment — remember the augmentation idea: we want to enrich the data, improve the quality, add something. At the very beginning of the course, when we talked about feature engineering, the same idea appeared: we enhance the features, we add features.

Data enrichment is the augmentation idea — improving quality by adding something, not just moving what exists. Feature engineering is the same mindset applied at the model level: you enhance the features, you add features. Both are about making the raw material more valuable before it is used.

Integration tools are also becoming AI-enabled, which changes who can use them and for what.

Q: How do integration tools support AI projects specifically? A: AI-enabled tools help you create queries — remember the MongoDB demo — help you build a chatbot, and help you build recommendation systems. That is what makes integration tools useful for AI projects.

Finally, the reason these tools can reach so many different systems is the connector — a pre-built adapter that speaks the target system's protocol.

Q: What feature allows integration tools to access diverse sources and targets seamlessly? A: Out-of-box and configurable connectors. Each kind of data comes with a different tool — streaming data, databases — so the integration tool should be customizable, ship with out-of-box capabilities, and have configurable connectors: database connectors, XML, JSON, API, web services. RapidMiner and Orange Miner let you connect to CSV files, Excel files, XML files, SQL files, and Python scripts.

A connector is like a power adapter for data: the same tool plugs into a database, a file, an API, or a web service, and the adapter handles the differences. The more connectors a tool ships with, the wider the range of sources and targets it can reach without custom code.

11.3 A Study Tool — NotebookLM for Exam Preparation

A large part of the class was a practical study tip: NotebookLM (notebooklm.google.com). The instructor uses it to train banking professionals in Malaysia and Singapore — for example, uploading FinTech policy documents to understand a framework better — and it worked very well. The idea is simple: feed the tool your own material, and let it turn that material into study aids.

11.3.1 The NotebookLM Workflow

Log in with your Google account, go to notebooklm.google.com, and create your own notebook. Then add sources: you can upload local files — for example, a course session document on data profiling and validation. The key behavior: NotebookLM builds its own corpus, its own vector database, its own text corpus from the sources you add, and it will only refer to those local sources — it will not go and look things up on Google. If you upload a document about data profiling and validation, the answers come only from that document.

A vector database is a storage system that holds text as lists of numbers (vectors) so that similar meanings sit close together; NotebookLM uses it to match your question against the parts of your sources that are most relevant. This grounding behavior is what makes the tool useful for study: the answers are constrained to what you gave it, so they reflect your course material rather than general internet content.

You can also create a prompt — this is prompt engineering applied to your own material. You write exactly what you want in plain words, and you can also add more sources with a prompt. Everything you generate is constrained to the available local sources.

11.3.2 What NotebookLM Can Generate

Once sources are loaded, NotebookLM can generate:

  • Flash cards — a quick way to review a topic.
  • Mind maps — ideal for complex topics with many architecture diagrams. The instructor showed a mind map of a framework for data schema integration mapping: direct mapping (mix the tables with minimal or no transformation), transformed mapping (use the existing structure, require transformation, do a split or merge), and custom data mapping (very client-specific data).
  • Audio — you can generate audio and download it, then listen while traveling by train or bus.
  • Quizzes, data tables, and generated presentations from the same sources.
  • Study guides and blog posts, and you can generate reports from an Excel file.

The instructor showed a mind map for the data profiling and validation topic: ML data readiness requires proper data profiling, a proper validation framework, and skewing and drifting detection; data profiling covers missing value identification and unexpected feature detection; skew and drift covers drift detection metrics and data leakage; the common data problems include technical noise (formatting loss, random corruption, audio background noise), statistical challenges, and labeling and logical problems.

The reasoning behind the tip: our brain is very powerful at registering images, so mind maps and flash cards are a fast way to prepare for an exam — instead of going through 90 to 95 pages of content, you quickly look at the flash card, which connects the topic for better exam understanding and writing. The instructor also used the tool for teaching software engineering — the spiral model diagram made the concept easy for students to understand.

11.3.3 Sources Beyond Local Files

NotebookLM is not limited to uploaded documents. You can point it at web sources and even at code repositories.

Q: Could I pass a GitHub repository here and get a mind map of that code? A: Yes, you can — you can even understand the code. Add the repository link as a web source, then ask questions like "tell me what it has and what versions are available" and it goes through the repository and gives you the answer quickly, instead of you skimming the entire thing.

You can add web sources directly — for example, a GitHub file or a documentation link. You can also add an Excel file: read it, make a decision, generate reports, create a study guide or a blog post. For prompt engineering, you can even ask a general-purpose tool to write a prompt for NotebookLM: "I am going to use NotebookLM for creating a mind map or reading a document that illustrates the mapping and architecture — it should only provide the explanation using the local source" — then take that prompt, add your exclusions, and run it in the notebook.

Q: Should the information be unauthenticated for this to work? A: For public web sources it reads directly from the link. For organizational data, see the safety rule below — you must control what you feed it.

11.3.4 Using NotebookLM Safely

The tool is powerful, but it is a cloud service, and that raises the question of what you are allowed to put into it.

Q: Is it safe to use NotebookLM with organizational data or GitHub links? A: Use it with a pinch of salt. Organizational data and private repositories may need to pass through an authentication layer. The safe pattern is to build your own local corpus — a local LLM — add only the sources you want, delete everything you do not want to expose, and keep full control of sources and prompts. Never put company data you are not allowed to share.

The instructor's rule of thumb: you have control over what goes into the notebook — do not put everything; only give what you want, and exclude what you do not want. A guide for generating flash cards and mind maps with NotebookLM was prepared and shared with students.

Scope: NotebookLM is grounded in your sources, but grounded is not the same as private. Anything you upload leaves your machine and lives in a cloud notebook, so the safety rule is: treat the notebook like a public whiteboard. Organizational data and private repositories may need to pass through an authentication layer before they are allowed in; when in doubt, build your own local corpus with a local LLM and keep full control of sources and prompts. Never expose company data you are not allowed to share.

11.3.5 Exam Strategy with NotebookLM

Exam note: For exam preparation, upload the course material — even complex topics with many architecture diagrams — then generate flash cards and mind maps to understand the concepts quickly. Generate audio to review while you commute. Instead of going through 90 to 95 pages of content, review the flash card: it connects the topic for better exam understanding and writing. This is an easy, low-effort study tip that works for any subject.

11.4 Machine Learning Experimentation

11.4.1 From Clean Data to a Trained Model

You have heard about the data rule: 70-30, 80-20, 50-50, 60-40 — different ways to split data for model training. Any machine learning model needs to be trained on a huge volume of data to identify patterns. The preparation flow, revisiting earlier classes: first, design the problem statement. Second, have access to the data set, clean and clear — the data engineering is already done: data ingestion, data digestion, data massaging, data scrubbing — garbage in, garbage out. This builds on the CRISP-DM process (business understanding, data understanding, data preparation, modeling, evaluation, deployment), feature engineering, and data pipelining studied earlier. Once the data is clean, decide which algorithm will be used for the model — random forest, LSTM, an ensemble method — and which parameters and hyperparameters you need. Then split the data set into training and testing sets, prepare the model algorithms for training, and train.

Two kinds of numbers matter here, and students often mix them up. Parameters are the values the model learns from the data during training — the weights of a regression, the split points of a tree. Hyperparameters are the values you choose before training starts — the learning rate, the number of trees, the batch size. You set the hyperparameters; the data sets the parameters.

Exam note: in project and dissertation reviews — the instructor has examined data science projects for over a decade — the first question is always: what are the parameters you are considering? Performance parameters and hyperparameters: what was the learning rate? What was the accuracy? What was the error rate? Based on your problem statement you must come up with the critical parameters, and use those parameters to identify how good your proposed model is — whether it is a classic ML model, deep learning, an ensemble, or a Gen AI agent. Define them early, and be ready to justify them.

11.4.2 Training, Validation, and Test Data

Training data is the data we use to train the model — we make adjustments to the functions, to the machine learning algorithm, and ensure the objectives are met. Validation data is used during the training process. Test data comes last: after training and validating the system, we take some 10 to 20% of the data that was not used before, to test the model's accuracy, error rate, and learning rate.

The three sets play three different roles:

Set When it is used What it is for
Training During training The model learns: weights, splits, and other parameters are adjusted on this data
Validation During training, alongside it Guides the training: tune hyperparameters, watch for overfitting, pick checkpoints
Test After training ends Final, untouched measure of accuracy, error rate, and learning rate

Together, these three sets let us understand underfitting and overfitting, any model bias, and any drift between the data sample and the model.

Normally we split the data into a training set and a test set, for example an 80-20 rule, or 70-30, or 60-40. You try different splits, run the tests, and come up with the results. You can also go with K-fold cross-validation to build a strong, reliable model.

11.4.3 Cross-Validation — K-Fold, LOOCV, and LPOCV

K-fold cross-validation: split the data into \(K\) equal parts, called folds:

\[ D = D_1 \cup D_2 \cup \cdots \cup D_K \]

where \(D\) is the whole data set and \(D_1, \ldots, D_K\) are the \(K\) folds. Designate one fold as the holdout — the part we keep aside. Train the model on the other \(K - 1\) folds, then test the model on the holdout fold. Repeat the process \(K\) times, each time selecting a different fold as the holdout — a kind of randomization. After \(K\) rounds, every record has been tested exactly once. You get a cross-validated score by averaging the \(K\) round scores:

\[ E_{\text{CV}} = \frac{1}{K}\sum_{k=1}^{K} E_k \]

where \(E_k\) is the error measured on the holdout fold \(D_k\) in round \(k\), and \(E_{\text{CV}}\) is the final cross-validated error. The average turns \(K\) separate results into one number you can compare across configurations.

The instructor's picture: take a big chocolate and cut it into ten parts. Keep one part for the final tasting, and use the other nine for the training.

Worked example — ten-fold on 100 records. Suppose you have \(N = 100\) records and you choose \(K = 10\). Each fold holds:

\[ \frac{N}{K} = \frac{100}{10} = 10 \text{ records} \]

  • Round 1: test on records 1–10, train on records 11–100 (90 records).
  • Round 2: test on records 11–20, train on records 1–10 and 21–100 (90 records).
  • Round 3: test on records 21–30, train on the remaining 90 records.
  • … and so on, until…
  • Round 10: test on records 91–100, train on records 1–90 (90 records).

Every one of the 100 records is used for testing exactly once and for training exactly nine times. Suppose the ten round error rates are 0.12, 0.10, 0.14, 0.11, 0.13, 0.09, 0.15, 0.10, 0.12, 0.11. Then:

\[ E_{\text{CV}} = \frac{0.12 + 0.10 + 0.14 + 0.11 + 0.13 + 0.09 + 0.15 + 0.10 + 0.12 + 0.11}{10} = \frac{1.17}{10} = 0.117 \]

Final answer: the cross-validated error is 0.117 (about 11.7%). Sense-check: each round error lies between 0.09 and 0.15, and the average 0.117 sits comfortably inside that range — no single fold distorted the result, which is exactly what averaging across folds is supposed to guarantee.

The K in K-fold is your choice: three-fold, four-fold, five-fold, and so on. Finally you aggregate the results and decide your training and testing data sets. Why do this instead of a single split? Because the partition can be a bit blunt and you can ignore some important information; with K subsets instead of two parts, you have more luxury of control, and then you can do the validation properly.

LOOCV — leave-one-out cross-validation — simply leaves one data observation out of the training, which is where the name comes from. It is a special case of K-fold in which \(K\) equals the number of observations \(N\):

\[ K = N, \qquad \text{round } i \text{ trains on } N - 1 \text{ observations and tests on observation } i \]

In this situation we end up using all the available data in training, which lowers the bias of the model. The execution is longer, because we have to repeat the process \(K\) times — for \(N = 100\) records, that is 100 separate training runs. The instructor's analogy: like a movie with a surprise element that no one can guess — not even from the trailer — the twist appears only at the end; that surprise element is the single observation left out.

LPOCV — leave-P-out cross-validation — instead of leaving one single observation out, we assess the model against a group of \(P\) observations. For a data set with \(N\) observations, we use \(N - P\) data points for training, and the remaining \(P\) for testing:

\[ \text{train on } N - P \text{ points}, \qquad \text{test on } P \text{ points}, \qquad \binom{N}{P} = \frac{N!}{P!\,(N - P)!} \text{ possible test sets} \]

where \(\binom{N}{P}\) counts how many different groups of \(P\) observations can be chosen out of \(N\). The number of samples created for LPOCV grows, so it is computational; you can combine multiple samples. For \(N = 100\) and \(P = 10\), that is \(\binom{100}{10} \approx 1.73 \times 10^{13}\) — about seventeen trillion possible test sets — so in practice you take a handful of random groups instead of all of them. For example, you might look at one particular customer segment or one particular city as the sample of \(P\) observations. These ideas connect back to sample bias and how we sample data. Shuffle split: we can also shuffle the splits — use a random number — so the 80-20 split is not always the same 80%.

Q: Is K-fold cross-validation also data augmentation? A: Yes, in a sense. But apart from the folding itself, within each fold we may change something: for example, change the order of the data — the entropy order of the data — or change one or two features, doing some kind of alteration between experiments. The model stays the same while the data changes.

The answer connects to the four experiment types below: when the model stays the same and the data changes, you are running a data-augmentation experiment — each fold variant is a slightly different view of the same problem.

Q: Could you walk through LOOCV? A: We leave one data observation out of the training — we are not using it at all. It is a special case of K-fold in which K equals the number of observations, so we end up using all the available data in training, which lowers the bias of the model. The cost: execution is longer because we repeat the process K times.

Q: And LPOCV? A: Instead of leaving one single data observation out, we assess the model against a group of P observations. For a data set with N observations, we use N minus P data points for training, and the remaining P for testing — for example, one particular customer segment or city as the sample.

11.4.4 What Counts as an ML Experiment

Engineering means a structured way of doing things. An ML experiment is the engineering process of testing a machine learning model in that structured way: every experiment runs a model over a data set, makes a prediction, validates the prediction, and notes it down.

The instructor's analogy from electronics: when you test a pressure sensor, you give some input pressure, you use the sensor, and you measure the output. Testing gates, counters, registers, and microprocessors works the same way — you run the test again and again. Machine learning is the same: you give the data set through the data pipeline, you get training data, you apply the model, but in a very structured, organized way — what will be the input this time, what will be the input in the next cycle.

We do not just pick a model on first success — "random forest is working beautifully, linear regression is working beautifully, CNN with RNN is working beautifully" — no, we do not fix that model. We conduct many experiments before coming to a particular production model. Between experiments we make small changes in the model parameters and in the data configuration, and we track the changes and the results: which parameter did better, which parameter did poorly, and what we improved.

The full experiment loop: give the training data, run the machine learning algorithm, find out which parameters perform better and which perform worse, improve them, then validate with the validation data set; with the optimal hyperparameters settled, run the testing data set — the final data set for testing — and choose the best model. Based on the best model, the model goes into the system, is deployed, and a retrain frequency is planned. We test for a particular period — one day, one week, one month — with different training combinations and different epochs and batches (the neural network terms for these passes).

Pitfall: fixing the first model that works. "Random forest is working beautifully" is not a reason to stop — the first success may be a lucky split of the data. The disciplined pattern is: one variable at a time, measured and recorded. Never trust a single run; never deploy without a comparison.

11.4.5 Four Types of ML Experiments

There are four types of experiments we normally do: model selection, feature engineering, hyperparameter tuning, and data augmentation. Each type changes exactly one component of the pipeline while the others stay fixed, so the effect of the change can be isolated.

  • Model selection: the data set remains the same; you change the models between the experiments. Say you have a random forest, XGBoost, AdaBoost, or other classification models — you try them one by one on the same data set and then select.
  • Feature engineering: the model remains the same; you diversify and change the data features between experiments. Maybe you check with the age feature, then with another feature — feature-engineering-based experiments.
  • Hyperparameter tuning: the models remain the same; you change only the hyperparameters — tune the learning rate, tune the bias, tune certain parameters — and see whether the model performs better.
  • Data augmentation: the model remains the same; you change the data set between experiments — evolve the data, make adjustments to the data. One experiment gives only ages 5 to 16, the next gives ages 16 to 25 — you play with the data while the model stays the same.

Each of the four diversifies a different component of the pipeline across experiments. The table makes the pattern visible:

Experiment type What changes What stays the same Typical question
Model selection The model (random forest, XGBoost, AdaBoost, …) Data set Which algorithm wins on this data?
Feature engineering The features (age, another attribute, derived columns) Model Which features help the most?
Hyperparameter tuning The hyperparameters (learning rate, bias, …) Model What setting performs better?
Data augmentation The data set (ages 5–16, then 16–25, …) Model How does more/varied data change results?

11.4.6 Model Analysis and Validation

After the experiments you analyze and validate the model: in terms of accuracy, in terms of how strong and reliable it is, and how well it generalizes. We use statistical measures — check the metrics like accuracy, precision, use a confusion matrix, and find the F1 score. A confusion matrix is a table that compares predicted labels against true labels: true positives, true negatives, false positives, and false negatives. Precision asks "when the model said positive, was it right?"; recall asks "of all the real positives, how many did the model find?"; the F1 score is their combined single number. This model validation applies to different industries — for example, credit risk for a customer.

The six-step model validation checklist, walked through with a credit-risk example:

  1. Conceptual soundness — does this model make sense? Is the model conceptually sound?
  2. Data quality — all the data quality parameters we checked: completeness, accuracy, the 7 or 8 data quality parameters studied at the very beginning. Can we trust the training set?
  3. Process verification — was the process properly followed from start to end, and was the model built and deployed correctly? This is verification: independent review, reproducible tests, lockdown change control. Did we do the right thing?
  4. Outcomes and analysis — what outcome do we expect, whether it is regression, classification, or clustering, and does the model perform and stay fair? This is the actual validation — in terms of what we really got.
  5. Ongoing monitoring — keep monitoring the model continuously, so that we do not have any drift.
  6. Governance — who is accountable? Are we doing the right thing for the organization? Who has clear roles, are we following the policies, who has the correct rights? Governance also means masking the data based on the domain — HIPAA compliance, PCI DSS compliance, or government regulation compliance.

Worked example — the checklist on a credit-risk model. A bank builds a model that scores whether a loan applicant will default.

  1. Conceptual soundness: does it make sense to predict default from income, debt-to-income ratio, and repayment history? Yes — these are established drivers of credit risk, and the model's logic agrees with what credit analysts expect.
  2. Data quality: completeness and accuracy checks on the training set — are income fields missing for 20% of applicants? Is the repayment history column correctly coded? If the training set cannot be trusted, nothing downstream can be.
  3. Process verification: was the model built and deployed correctly — independent review of the code, reproducible training runs, and lockdown change control so nobody silently swapped a model version?
  4. Outcomes and analysis: for a classification outcome (default / no default), does the model perform — accuracy, precision, F1 — and does it stay fair across age and income groups?
  5. Ongoing monitoring: watch the score distribution each month so that drift in the applicant population is caught early.
  6. Governance: who is accountable for this model's decisions? Clear roles, policies followed, correct access rights, and masking of sensitive customer data per HIPAA, PCI DSS, or government regulation.

Final answer: the bank walks the same six steps before the model touches a single loan decision. Sense-check: the checklist moves from "does it make sense" (step 1) to "can we trust the data" (step 2) to "was it built right" (step 3) to "does it perform" (step 4) to "does it stay healthy" (step 5) to "who answers for it" (step 6) — a natural order from idea to accountability.

Q: Would you walk through the six-step model validation checklist with a credit-risk example? A: Step one, conceptual soundness — does this model make sense? Step two, data quality — all the data quality parameters we checked (completeness, accuracy); can we trust the training set? Step three, process verification — was it built and deployed correctly (independent review, reproducible tests, lockdown change control)? Step four, outcomes and analysis — does it perform and stay fair; this is the validation of what we really got. Step five, ongoing monitoring — keep monitoring so there is no drift. Step six, governance — who is accountable, clear roles, policies, correct rights, and masking the data per domain: HIPAA, PCI DSS, or government regulations.

11.4.7 Quality Control vs Quality Assurance

The discussion on process verification raised the quality control versus quality assurance distinction, a classic confusion point. A student guessed that the quality assurance team has more people, because "they need to execute and see whether it works." The roles were flipped: quality control does the majority of the work in traditional organizations and has more people — QC executes the tests, validates everything, and puts the stamp of authority on a release; it is reactive. Quality assurance is very small (often one or two SQA people): QA provides the processes and checks "are we doing the right thing?" — it is verification, an independent review. In a typical company you will find many testers — manual testing, automated testing — but very few QA people.

Q: Which team has more people — quality control or quality assurance? A: Quality control is predominantly bigger in any organization: they are the ones actually validating, executing, and checking everything — they put the stamp of authority, and they are reactive. Quality assurance provides the processes and checks whether we are doing the right thing — it is verification.

The modern stack flips the balance: with shift-left, everything follows automation. The moment you create test cases and plans, you can automate them — black-box validation runs in the code — so you need more quality assurance and governance people than validators. Some teams now remove dedicated QC entirely: developers write behavior-driven (BDD) and spec-driven tests while talking to product managers, so the QC role disappears — QC is hidden. The instructor gave the pattern from industry: a company used to have 70 developers and 20 to 30 testers; now there are 20 developers, and they do the testing themselves.

Q: Does not the modern stack flip the traditional balance? A: Yes. In the traditional world the inclination toward control is higher. In the modern stack, quality assurance outnumbers quality control because everything follows shift-left: once you create the test cases and plans, you automate them, and black-box validation happens in code. Some teams even remove dedicated QC people — developers write behavior-driven and spec-driven tests with the product managers, so QC is hidden.

Recap: QC executes and validates — it is reactive and, traditionally, big. QA defines processes and verifies — it is proactive and, traditionally, small. Shift-left automation is reversing the balance: when tests run in code, the assurance and governance mindset grows while dedicated validators shrink. Both roles still exist; their sizes are what changed.

11.5 Metadata — Data About Data

11.5.1 What Metadata Is

When you work on a laptop or any system, you have lots and lots of information: previous information about something, details of everything — customer data, product data. On top of that sits data about data. A picture of a cat is just data — the picture. Metadata is more information about it: more details, more granular information, detailed information for further analysis.

For a customer record, the questions that metadata answers are: why do we have "customer" as the entity? Why do we have "age" as an attribute? When was it actually collected? When was it modified? Who is the actual owner? Where did it come from — where did it originate, where is it stored? How do we collect the data — what was the method used to bring the data in, and how is it processed? All of that is metadata, sometimes called a data dictionary. Metadata empowers data discovery.

Hook: the picture is the data; the story around the picture is the metadata. A photo of a cat says "cat." Metadata says who took it, when, with which camera, at what location, and who is allowed to use it. The data answers "what"; metadata answers "why, when, who, where, and how."

A data dictionary (metadata in tabular form) is the same idea written down: one entry per field, with meaning, type, owner, and rules. When someone asks "why do we store age as a number and not a text string?", the answer lives in the metadata.

11.5.2 Why Metadata Matters

Q: Why do we need metadata — why do we need all these details? A: Four reasons. It empowers data discovery: a strong metadata layer acts as a search engine for your enterprise data, so you can find the relevant data sets without tribal knowledge of Slack threads. It increases data literacy: metadata bridges the gap between data producers — the engineers — and data consumers — the business users — and becomes the common language across departments. It enables trust and compliance in regulated industries like finance, healthcare, and pharma, where knowing who touched what data and when is non-negotiable. And it accelerates cloud modernization: cloud migrations fail not because of computer power, but because teams do not know which data is safe to move, delete, or refactor — metadata gives clarity.

Each of the four reasons deserves a closer look. A strong metadata layer is a search engine for enterprise data: you find the relevant data sets without tribal knowledge of Slack threads — you do not need to know which channel or which person holds the answer. Metadata increases data literacy because different people look at the same data differently — a doctor, an engineer, a psychologist, and a professor each see different views and dimensions of the same record. Metadata is a kind of common language across departments, similar to UML, the unified modeling language used for designing systems — metadata plays the same role for data. In regulated industries (finance, healthcare, pharma), knowing who updated the data and when is non-negotiable — did a doctor change the data, or a software engineer? Finally, cloud migrations fail not because of computer power but because teams do not know which data is safe to move, delete, or refactor; metadata gives the clarity to decide.

Scope: metadata is a language, and a language only works if both sides speak it. If the engineers write the metadata and the business users never see it — or the reverse — the bridge collapses. Discovery, literacy, trust, and migration all depend on metadata being written, kept current, and read by both producers and consumers.

11.5.3 The Five Dimensions of Metadata

Q: What are the five dimensions of metadata? A: Purpose — the intended use or reason for the data. Date and time — when the data was collected or modified. Ownership — who is responsible for the data. Methodology — how the data was gathered or processed. Location — where the data originates or is stored.

These dimensions are very clear when applied to a customer record: what is the purpose of this data? Why is customer the entity and age the attribute? When was it collected, when was it modified, who is the owner, where did it come from and where is it stored, and by what method was it collected and processed? This is how we use metadata in our day-to-day life.

Dimension Question it answers On a customer record
Purpose Why does this data exist? The entity "customer" exists so the business can track its clients; "age" exists as an attribute for analysis
Date and time When was it created or changed? Collected 18 June 2025, modified 19 June 2026
Ownership Who is responsible? The customer-data team
Methodology How was it gathered or processed? Collected through the sign-up form, validated by the intake pipeline
Location Where does it originate or live? Originated in the CRM, stored in the data warehouse

11.5.4 Types of Metadata

Q: What are the types of metadata? A: Descriptive metadata identifies and provides information about the data for easy discovery. Structural metadata describes the structure of the data — the parent-child relationships and what sits in between. Administrative metadata covers who has access: security permissions and compliance. Technical metadata covers specific technical details — compatibility, version, format, extension, which browser you can use. Statistical metadata records the sampling method used, how the data was sourced, and the different steps.

As the instructor put it: computer science is common sense, and all of computer science is derived from other fields — from the names themselves you can understand the types. Descriptive means descriptions of the data; structural means the relationships — what is the parent, what is the child, what is in between; administrative means who has access, security permissions, compliance; technical means the technical details — compatibility, version, format, extension, which browser; statistical means the sampling method, how we source the data, the steps and the method.

Type Plain meaning Example
Descriptive Identifies and describes the data for easy discovery Title, author, subject of a data set
Structural The shape of the data — parent-child relationships A table's foreign keys, the hierarchy between tables
Administrative Who has access — security and compliance Access rights, who may read or edit
Technical Technical details of the files Format, extension, version, compatibility, required browser
Statistical How the data was sampled and sourced Sampling method, source, processing steps

Recap: metadata is data about data — the granular detail that answers why, when, who, where, and how. Five dimensions (purpose, date and time, ownership, methodology, location) frame the questions, and five types (descriptive, structural, administrative, technical, statistical) organize the answers. The names tell you the meaning — computer science is common sense.

11.6 Metadata Categories for Machine Learning Data

When it comes to machine learning data, metadata falls into six categories: artifact, data set, feature, label, ML model, and pipeline. Each category answers a different question about the ML project, and together they cover the whole lifecycle — from the raw files in storage to the trained model in production.

11.6.1 Artifact Metadata

Artifacts are anything used for the project apart from the experiments themselves — any input or output of the runs. In software engineering, an artifact means anything like data, a document, a table, an SRS, or a test case. In the machine learning perspective, an artifact can be a data set, a model, the prediction results, or any other file — an Excel file, a CSV file.

We want to log all artifacts: where each one was referred from — an AWS S3 bucket, an NFS file system, a local file system — what version it has, and a preview so you can see what the artifact is like: run head, or head -5, to see at least five lines of the data. When you click that CSV file or that AWS source, you know what it is all about. An SRS (software requirements specification) is a document that states what a system must do; in ML terms, the artifact log plays a similar role for the files of the project — provenance, version, and a peek inside.

11.6.2 Dataset Metadata

Dataset metadata is the metadata above the data set: where it has come from, the locations of the unprocessed data, who is the responsible person and team that takes care of the data, what version stands behind the data's creation and updates, how it was updated, dataset restrictions, and the licensing part — which is very important. Licensing matters because a model trained on data you do not have the right to use is a legal risk: the metadata records what you may do with the data, and the model inherits that question.

11.6.3 Feature Metadata

Feature metadata is about a particular feature — and a feature is a variable: a reference to code, a reference to anything. What the feature reads, how it is processed — maybe a particular column or a particular value — how it is read, how it is updated, any restrictions, and the authorship: who is the author of the feature definition, when was the feature created.

This matters especially for multi-modal data. Take a bio record: it has a photo, it has many headings — maybe a LinkedIn page. Feature metadata answers: does everyone have access, or only some people? Is the link working, or is it a profile problem? What was the creation date? What are the restrictions on who can use it? A feature here is one input variable the model reads — for a customer record, features could be age, income, and the photo attached to the profile; each one gets its own metadata row.

11.6.4 Label Metadata

When we give a specific label to the features, we get label metadata — similar to feature metadata, like a lookup. We call the columns features because in neural networks — CNN, RNN — we do not store labels; we just use the features and do a lot of feature-related calculations and updates. But in traditional machine learning we use label metadata: what was the label source? What was the label confidence — was it a good label? When and how did the version change? When was the label name changed? A label is the answer the model is trained to predict — spam or not spam, default or no default; label metadata records where that answer came from and how sure the annotator was.

11.6.5 ML Model Metadata

Q: What does ML model metadata contain? A: The training parameters, the evaluation metrics, prediction examples, data set versions, testing pipeline outputs, and references to the model weight files — plus who created the model, which department, which code was used for training, and every version detail. That is where the pickle file comes in.

ML model metadata captures: what training parameters were used, what the particular version of the model is, how the model was built, who created it — which department — which parameters and which code were used for training. All the details about the machine learning model can be stored in a pickle file, and from the pickle file itself you can build the metadata. GitHub is a classic example of how this works: whatever machine learning code, what input features were used, when they tried, who updated, how they performed the evaluation — everything is visible. A pickle file is Python's standard serialization format — it saves an object (here, the trained model) to disk as a binary file so it can be loaded back later.

When models matter to someone — whenever they are touched, whenever they drive decisions — we need to know more about them. During ML experimentation and training we need to visualize the model and the data: the data version, the environment configuration, the code version, the commit version, the hyperparameters, the training metrics, the losses. We also look at hardware metrics — CPU, GPU, TPU, memory — how much memory the model was consuming — plus F1 and F2 scores, accuracy, ROC curve, confusion matrix, and the model's predictions. An ROC curve (receiver operating characteristic) plots the true-positive rate against the false-positive rate at every threshold, and the area under it is a threshold-free score of how well the model separates classes.

Worked example — the pickle file as a metadata carrier. A data scientist trains a random forest classifier and saves it with pickle.dump(). The file itself is binary, but it carries metadata with it: what the source was, the data set details, the algorithm (for example, random forest classifier), when it was created — and you can load the model and rerun it. Later, anyone with the pickle file can answer: which model is this? Random forest classifier. On what data was it built? The customer data set, version 3. When? 18 June 2025. The metadata needed to identify, audit, and reproduce the model travels inside the artifact itself.

Trained models keep collecting metadata as you train: when the training happened, the model package, what the model binary is, where it was stored, the version number, which data set that model version uses, the evaluation record, and the experiment version. For version control, the pickle file itself carries the details — what the source was, the data set details, the algorithm (for example, random forest classifier), when it was created — and you can load the model and rerun it. Git plus DVC (data version control) and MLflow give comprehensive ML versioning: version control of code, data set, hyperparameters, and environment; references to downstream data sets and models; drift metrics — data drift, concept drift, performance drift — and hardware metrics. DVC (Data Version Control) treats data sets like Git treats code — versioned, diffable, and recoverable; MLflow tracks experiments — parameters, metrics, artifacts, and model registries — so every run can be compared with every other run.

11.6.6 Pipeline Metadata

Pipeline metadata covers the whole journey: from the sources — different types of sources, from legacy proprietary systems — how we got the data, which database was referenced, when a new branch came in, and the complete CI/CD workflow. This is the DAG — the directed acyclic graph — from the previous class. A DAG (directed acyclic graph) is the standard picture of a pipeline: nodes are tasks, arrows are dependencies, and the graph has no loops, so every task runs after its inputs are ready. How the data was collected, how it was finally stored in the target system: what the inputs were, what the output steps were, whether there are cache steps in between, whether there are staging steps, which pipeline runs they came from, what binaries were used — there may be many in-between programs running validation checks. All those details are pipeline metadata.

Recap: six categories of ML metadata — artifact (files and where they live), data set (origin, version, licensing), feature (definition, authorship, access), label (source, confidence, version), ML model (parameters, metrics, pickle file, Git + DVC + MLflow), and pipeline (sources, CI/CD, the DAG, cache and staging steps). Together they make every stage of the ML project traceable.

11.7 The ML Metadata Store — Why and How

11.7.1 Metadata Needs a Dedicated System

Metadata has many dimensions and lots of information — about the data, about the features, about the ML model, about the pipeline: who has access, versions, everything. That means the metadata you use for your machine learning or deep learning project requires a database, a repository, a dedicated system or server, or a dedicated storage space of its own. Metadata storage is not an afterthought; it is part of the project design. You would not run a project without a place for the code; the metadata deserves the same planning.

11.7.2 A Relational Database Example

Relational databases already do this. In Oracle, Informix, or Sybase, there is a separate metadata repository containing many tables. For example, there is a table called user_objects. This user_objects table is metadata, and it is stored in a separate tablespace — the system tablespace, also called sys tablespace or sys aux tablespace — not mixed up with user data, which lives in the user tablespace. Because metadata is so critical, it gets its own dedicated storage; this metadata dictionary is called the data dictionary. A tablespace is a unit of storage inside an Oracle-style database; putting metadata in its own tablespace means the database's self-description is never tangled with the data it describes.

The user_objects metadata table consists of many columns: the owner of the object, the object name, the sub-object name, the object type (is it a table or an index), the created timestamp, the last DDL timestamp, status, and more — lots of information about each particular object. DDL (data definition language) is the SQL for creating and changing structures — CREATE TABLE, ALTER INDEX — so the last DDL timestamp records the moment the object's structure last changed.

11.7.3 A Worked Example — Tracking One Feature

You can build your own metadata table — an Excel sheet, a metadata database, whatever fits. Columns you might track: the feature or attribute name, the ML model ID that uses it, when it was originally created, which data set it belongs to, the last modified date, and a modified comments field.

Worked example — one row for one feature. The instructor showed a live row:

Column Value
Feature name customer ID
ML model that uses it classifier model defined in YAML
Originally created 18 June 2025
Belongs to data set customer.xlsx
Last modified 19 June 2026
Modified comments modified for expansion — name changed from "cus ID" to "customer ID"

That row answers: which model consumes this feature? A classifier defined in YAML (a plain-text configuration format). When was it created? 18 June 2025. Where does it live? In the data set customer.xlsx. When did it change and why? 19 June 2026, renamed from "cus ID" to "customer ID" so the name could support more of the business. That is metadata — and you can add more and more fields onto it.

Q: What should a metadata system minimally keep track of? A: In the case of features and tables, it should minimally keep track of the feature definitions, which versions are used in the model definitions, and which trained models use them. A simple example is the Excel row for the feature "customer ID": which ML model uses it (a classifier defined in YAML), when it was originally created (18 June 2025), which data set it belongs to (customer.xlsx), when it was last modified (19 June 2026), and why — renamed from "cus ID" to "customer ID" for expansion.

Most organizations start building their data science projects without a solid metadata system, and they regret it later. Always build the metadata system — whether you use ServiceNow or any other tool, they all have built-in metadata, so there is no excuse to start without one.

Pitfall: starting without metadata. The pattern repeats across companies: the data science project begins fast, nobody records what the features mean or which model uses them, and six months later nobody can answer "what does this model depend on?" or "who changed this column and why?" — the exact questions the one-row example answers. The fix is cheap at the start and expensive later: build the metadata system from day one.

11.7.4 One System or Many

Two ways to structure the metadata system. Option one: build one system that tracks all the metadata from all the sources — one metadata system, one metadata table, one metadata database. The drawback: if that one data system fails, we are into soup — a single point of failure.

Option two: build multiple separate systems, one for each task — one metadata system for the training data, one for the ML pipeline, one for the model. This gives you decoupling. The drawback: more things to maintain. Both are valid; you trade simplicity against isolation.

Option Benefit Cost
One system for everything Simple, one place to look Single point of failure
Several systems by task Decoupling — one failure does not take everything down More to maintain and keep in sync

11.7.5 Build or Buy

Building your own metadata system has advantages: no license cost, you can create it exactly for your use case, and you can improve and modify it as you go. The cons: you have to implement everything yourself, and you have to set up and maintain your own infrastructure. People who work in software always face these choices in their careers: understand the problem you are trying to solve, what tools you are going to use, and how much time it takes to build. With vector databases and LLM databases, you can also create more metadata inside them — build a corpus.

11.7.6 The ML Metadata Store

The ML metadata store is a centralized place to manage all metadata about experiments — including experiment logs, artifacts, models, and pipelines. It holds everything: your training data, the performance details of each run, what time it ran, how it ran, the artifacts you collected, which data set was used. It has a user interface for reading and writing model-related data, and you can have different clients: a metadata store, a metadata dashboard, and a metadata database. From the dashboard you can access all the model information, all the data set information, all the artifacts, and all the pipelines.

Q: Is the ML metadata store the same as data warehouse metadata or RDBMS metadata? A: No — do not confuse them. There is metadata in the data warehouse, metadata in the data lake, metadata in traditional RDBMS systems, and metadata in real-time systems; those already maintain their own metadata. We are not competing with that. For our ML projects we need a dedicated metadata store because our line of business is different. If you depend purely on Oracle, DB2, or SQL Server you may reuse that existing metadata database, but it may still not be optimal — the practical route is some existing metadata plus your own.

The dashboard logs all the updates, stores everything, displays it, and compares between different models and different experiments — experiment one, experiment two. This is how you answer the manager's question: last month you ran this project, this month you ran that project — what is the percentage improvement? You must be able to compare and give the results.

Recap: metadata is too important to share a drawer with user data — relational databases gave it its own tablespace, and ML projects give it a dedicated store. One system or many is a simplicity-versus-isolation trade; build or buy is a cost-versus-control trade. The ML metadata store centralizes experiments, artifacts, models, and pipelines behind a dashboard, so "what improved between last month and this month?" is a query, not a guess.

11.8 Metadata Repository, Registry, and Store

There are three ways to organize the metadata: a metadata repository, a metadata registry, or a metadata store. The instructor flagged this comparison as exam-important in advance — "put the word important." The three words sound similar, but they answer three different questions: where do objects live, which versions matter, and where do people go to find things?

11.8.1 Metadata Repository

Q: What is a metadata repository? A: A place where metadata objects are stored along with all the relationships between those objects. You can use GitHub as a repository: save all the evaluation metrics as a file created via a post-commit hook, or log parameters and losses during training to an experiment tracking tool — which in this context is actually a repository.

A concrete walkthrough: a public GitHub repository containing random-forest code. The repository files show what data was used, what was done, when the files were touched, and what the training data was. The metadata usually includes that view file — you can read it, see when it was created, see the history, and see the author. One such repository was created by a third-year computer applications student with a passion for machine learning, deep learning, and computer vision; nobody changed it afterwards. That is metadata as a repository. A post-commit hook is a script that runs automatically after every commit — the perfect place to append the latest evaluation metrics to a tracked file, so the repository keeps a running history of the model's performance.

11.8.2 Metadata Registry

Q: What is a metadata registry? A: A place where you checkpoint the important metadata — you register something you care about and want to easily find and access later. A registry is always for something specific; there are no general registries. For example, you may have a model registry that lists all the production models, with references to the ML metadata repository where the actual model-related metadata lives.

The bill-of-materials analogy helps: whenever we build a product and ship a version, we always have a readme file with it. One registry file says what this model is about, when it was made, and what the related data are. It is just one registry file associated with the artifact. A model registry is the curated list of the models that actually made it to production — the ones that matter — pointing back into the repository where the full detail lives.

Scope: a registry is not a warehouse. It checkpoints the few versions that matter (production models, approved data sets), while the repository holds everything — experiments included. If you put every draft into the registry, you lose the point of the registry: finding the important thing fast.

11.8.3 Metadata Store

Q: How is a metadata store different from a repository? A: A repository stores objects and their relationships; a store is a place where you go shopping for metadata for ML models. It is a central place where you find all the models, experiments, artifacts, and pipeline metadata. You come to the shop, search the metadata products, compare them, and pick — it is more a store than a repository.

The online-store analogy: like Amazon, Flipkart, or any big marketplace — everything is available; you search, and you find all the details. The repository asks "where does this live?"; the store asks "what do I want, and which one fits?"

11.8.4 Choosing Among the Three

You can go three ways: metadata repository, metadata registry, or metadata store. The choice changes what the store emphasizes. A metadata store can treat the pipeline as a first-class object: the metadata contains the pipeline as a first-class object, and you associate — for this pipeline, which model will I use, which data set will I use, which experiments can be done using this pipeline. Or you can treat the model and the experiments as first-class objects, depending on what you put in the center; the ML metadata store will do different things depending on that choice. A pipeline-first design and a model-first design are two examples of how this manifests. First-class means the thing is treated as a real entity of the system — it gets its own records, its own relationships, and its own pages in the store — not just a string in someone's log.

Decision When it fits
Repository You need a home for all artifacts and relationships — the full history of everything
Registry You need the few important versions flagged and easy to find — production models, approved data sets
Store Teams need to search, compare, and pick — models, experiments, artifacts, pipelines side by side
Pipeline-first store The pipeline is the unit of work — ask "which model and data set does this pipeline use?"
Model-first store The model is the unit of work — ask "which experiments produced this model?"

Exam note: questions on ML experiments and ML metadata are expected from this class and the previous class — a lot of good questions may come. A typical question gives a case-study scenario and asks: what do you recommend — a repository, a registry, or a store — and how would you go about it? Metadata is also a unified concept across many subjects — database subjects, business subjects, data visualization, Power BI, machine learning pipelines, data subjects — so this material pays off in several courses. When you answer such a scenario, name the option, say what it stores and what it emphasizes, and walk through how you would set it up.

11.9 Metadata Management and Lineage Tracking

11.9.1 Metadata Management

Metadata management is the process of capturing the metadata about everything happening in the machine learning models, then organizing it and maintaining it. It is like database management — creating a database, organizing it, managing it — or file management: capturing the file, organizing the file, maintaining the file. The pattern is the same in all three: capture, organize, maintain.

The types of metadata management follow the metadata types we already saw: if you care about descriptive metadata, it is descriptive metadata management; if it is administrative, administrative metadata management; if structural, structural metadata management. You manage the same five types — descriptive, structural, administrative, technical, statistical — but the "management" part is the process: capture it, organize it, keep it current.

11.9.2 Data Lineage — Forward, Backward, and Horizontal

Data lineage is how the data is evolving, how the data is getting created, and how the data is currently present — the state of the data.

Q: What is data lineage and what does it capture? A: It captures the origins of the data set, how it moves over time, and what happens to it between its initial creation and its present state — where it started, when and how it was changed. Take a customer balance: what was the balance one month ago, what was it last week, what was it yesterday, what is it now — what happened in between? Lineage helps you track errors back to the source.

Worked example — the lineage of a customer balance. A customer's account balance, tracked over four snapshots:

Snapshot When Balance
One month ago 7 July 1,000
Last week 31 July 850
Yesterday 6 August 620
Now 7 August 1,050

What happened in between: the drop from 1,000 to 850 was a bill payment of 150; the drop from 850 to 620 was a purchase of 230; the jump from 620 to 1,050 was a salary credit of 430. Final answer: the present value 1,050 is the end of a chain — 1,000 − 150 − 230 + 430. Sense-check: 1,000 − 150 = 850, 850 − 230 = 620, 620 + 430 = 1,050 — each step reconciles. If a credit-risk model scores "now," backward lineage shows exactly which events built that number, so a wrong balance can be traced to the offending transaction instead of guessed at.

Q: What are the three types of lineage? A: Forward lineage shows how data moves from source to destination. Backward lineage reveals the origin of the data. Horizontal lineage illustrates the flow of data across systems and processes — it cuts across all the systems, from the legacy systems through the ML pipeline, through the data lake, through all the different components.

Tracking the lineage provides where the data came from, what the relation is between this data and that data, when the data changed, what transformation happened in between, and why it changed over time. When lineage — the data history, the state of data changes — is properly integrated with metadata management, we can understand the data transformation better, track the data properly, debug, and solve problems. The three directions answer three questions: forward asks "where does this go?", backward asks "where did this come from?", horizontal asks "how does this move across the whole system?"

11.9.3 Experiment Tracking

Experiment tracking answers practical questions: what was the previous experiment? How many times did this project run? What parameters were used? What different training data was used? What evaluation data was used? Keeping track of all of these is where experiment tracking comes into the picture.

Tracking the experiments also means tracking what scripts were used — which Python scripts, any batch files, any INI files, any configuration files. Whether you work with a JavaScript front end, a full-stack setup, or an ML pipeline, you have a lot of configuration files. People in the room had used PyTorch and PySpark. IDE platforms — like Google Antigravity — show you the configuration files and the start file of a project, so when you want to go back to the project, you can go back. The instructor showed an Antigravity project live and promised a fuller walkthrough of configuration files in a later session. An INI file is a simple text configuration format (sections and key-value pairs); the point of tracking it is that the experiment's exact settings — not just its results — are part of the record.

Recap: metadata management captures, organizes, and maintains metadata; lineage tracks how the data evolved — forward (source to destination), backward (origin), horizontal (across systems); experiment tracking records runs, scripts, and configuration files so "what did we try and what happened?" always has an answer. Lineage plus metadata management equals the ability to debug from effect back to cause.

Exam Guidance Summary

Consolidated exam guidance from the class:

  • Exam note: expect questions on ML experiments and ML metadata — from this class and the previous class. A lot of good questions may come, so review the experiment loop and the six metadata categories before the exam.
  • Exam note: be ready for case-study scenario questions: given a scenario, would you recommend a metadata repository, a metadata registry, or a metadata store — and how would you go about it? The repository/registry/store comparison was explicitly flagged as important, so be ready to state what each one stores and when each fits.
  • Exam note: know the four types of ML experiments (model selection, feature engineering, hyperparameter tuning, data augmentation) and be able to explain what changes between experiments in each type — one component changes, the rest stay fixed.
  • Exam note: know the cross-validation family — K-fold, LOOCV (K equals the number of observations), LPOCV (train on N minus P observations) — and why each exists: bias reduction, data usage, computational cost.
  • Exam note: in project and dissertation reviews, expect the first question to be about parameters and hyperparameters — learning rate, accuracy, error rate. Define the critical parameters for your problem statement and use them to justify your model.
  • Exam note: the six-step model validation checklist (conceptual soundness, data quality, process verification, outcomes and analysis, ongoing monitoring, governance) — a credit-risk scenario was the example; the checklist is also the answer for "how do we validate a model before deployment?"
  • Exam note: know the quality control vs quality assurance distinction: QC executes and validates (reactive); QA defines processes and verifies (proactive); and know how shift-left automation is changing the balance.
  • Exam note: use NotebookLM for exam preparation — upload the course material, generate mind maps and flash cards, generate audio for review while traveling, and create quizzes. Instead of going through 90 to 95 pages of content, review the flash card to connect topics for better exam understanding and writing.
  • Exam note: metadata is a unified concept across many subjects — database subjects, business subjects, data visualization, Power BI, machine learning pipelines — so learn it once and apply it everywhere. The same five dimensions and five types keep appearing in different courses.
  • Exam note: the flashcard recap of data integration, ingestion, and pipelines is the base for the metadata content; expect those ideas to be connected. Know the distinction: ingestion collects, integration curates.

Key Industry Applications

  • Real-world: MongoDB supports AI projects in data integration settings — creating queries, building chatbots, and building recommendation systems.
  • Real-world: RapidMiner and Orange Miner ship out-of-box and configurable connectors to CSV, Excel, XML, and SQL files, and to Python scripts — the same tools also let you connect to databases, JSON, APIs, and web services.
  • Real-world: NotebookLM (notebooklm.google.com) is used for training banking professionals in Malaysia and Singapore, for teaching (the spiral model in software engineering), and for exam preparation — with strict control of local sources for organizational data.
  • Real-world: GitHub serves as a metadata repository with post-commit hooks that log evaluation metrics; DVC (data version control) and MLflow provide comprehensive ML versioning — Git versions code, DVC versions data sets, MLflow tracks experiments and models.
  • Real-world: Oracle, Informix, and Sybase maintain separate metadata repositories (for example, the user_objects table in a system tablespace); DB2 and Microsoft Azure SQL Server are other metadata sources you can reuse for ML projects.
  • Real-world: AWS S3 buckets and NFS file systems are common artifact stores referenced in artifact metadata, with previews via head.
  • Real-world: finance, healthcare, and pharma rely on metadata for who-touched-what-when audits, with HIPAA, PCI DSS, and government regulation compliance requirements.
  • Real-world: credit-risk model validation in banking uses the six-step checklist: conceptual soundness, data quality, process verification, outcomes and analysis, ongoing monitoring, and governance.
  • Real-world: ServiceNow and similar enterprise tools ship built-in metadata, so building a metadata system from scratch is rarely necessary — configure what exists, extend where needed.
  • Real-world: multi-modal data — like LinkedIn profiles with photos and headings — illustrates feature metadata: access, creation date, and restrictions.
  • Real-world: Google Antigravity is an IDE platform that keeps configuration files and start files so ML projects can be resumed later.
  • Real-world: Agile, DevOps, and ML Ops with CI/CD pipelines drive continuous data and the demand for always-on self-service data.

DMML Lecture 11 notes · Machine Learning Experimentation and Metadata

Data Management for Machine Learning· postgraduate· 2026-08-07

Sections Breakdown

1Opening Story — Attachment, Mindset, and Ownership

The burning-house story and the cockroach theory: attachment, ownership, and responding instead of reacting.

2Flashcard Recap — Data Integration, Ingestion, and Pipelines

Flashcard recap: data integration, ingestion vs integration, pipelines, real-time data, and integration tools.

3A Study Tool — NotebookLM for Exam Preparation

NotebookLM as a study tool: local corpus, flash cards, mind maps, audio, and safe use with organizational data.

4Machine Learning Experimentation

The experimentation loop: train/validation/test data, K-fold, LOOCV and LPOCV, four experiment types, model validation, QC vs QA.

5Metadata — Data About Data

Metadata as data about data: why it matters, the five dimensions, and the five types.

6Metadata Categories for Machine Learning Data

Six metadata categories for ML data: artifact, data set, feature, label, ML model, and pipeline.

7The ML Metadata Store — Why and How

Why metadata needs a dedicated system, the relational database example, and the ML metadata store.

8Metadata Repository, Registry, and Store

Metadata repository, registry, and store — what each stores and when each fits.

9Metadata Management and Lineage Tracking

Metadata management, data lineage (forward, backward, horizontal), and experiment tracking.

10Exam Guidance Summary

Consolidated exam guidance from the class, across all topics.

11Key Industry Applications

Real-world applications and tools mentioned during the class.

Postgraduate students in Machine Learning and Data Management

Exam Revision Notes

Below is the distilled, exam-ready core. Every entry comes from the full explanation above. Use this section for rapid review; return to the main notes when a point needs more context.

Opening Story — Attachment, Mindset, and Ownership

Must-know: Attachment to outcomes creates sorrow; ownership without emotional attachment is the healthy middle. Reactions are instinctive; responses are well thought out — respond, don't react.

Top pitfall: Treating the trigger (the cockroach, bad news, a failed experiment) as the cause of distress instead of your reaction to it.

Self-check: Why was the waiter not disturbed by the cockroach?

Connects to: Machine Learning Experimentation.

Flashcard Recap — Data Integration, Ingestion, and Pipelines

Must-know: Integration combines sources into one unified fact base with quality handled; ingestion just collects. Integration requires more domain expertise because it knows what applications need. Gartner's tasks: transformation, cleansing, loading and analytics, enrichment.

Top pitfall: Confusing ingestion with integration — ingestion collects everything without improving quality; integration curates through merging, filtering, and removing noise.

Self-check: Why do integration pipelines need more expertise than ingestion pipelines?

Connects to: Metadata — Data About Data, Metadata Management and Lineage Tracking.

A Study Tool — NotebookLM for Exam Preparation

Must-know: NotebookLM answers only from its local corpus (vector database). Generate flash cards, mind maps, audio, quizzes. For organizational data: control sources and prompts, never expose company data.

Top pitfall: Uploading company data or private repositories to a cloud notebook without an authentication layer — grounded does not mean private.

Self-check: Does NotebookLM look up answers on Google?

Connects to: Machine Learning Experimentation.

Machine Learning Experimentation

Must-know: K-fold: split into K folds, train on K-1, test on the holdout, repeat K times, average the errors. LOOCV: K=N, lower bias, longer execution. LPOCV: train on N-P, test on P. Four experiment types change one component at a time. Six-step checklist: conceptual soundness, data quality, process verification, outcomes and analysis, ongoing monitoring, governance. QC executes and validates (reactive); QA defines processes and verifies.

\[E_{\text{CV}} = \frac{1}{K}\sum_{k=1}^{K} E_k, \qquad K = N \text{ for LOOCV}, \qquad \binom{N}{P} = \frac{N!}{P!\,(N-P)!} \text{ test sets for LPOCV}\]

Top pitfall: Fixing the first model that works; confusing QC (executes and validates) with QA (defines processes); forgetting that LOOCV costs N training runs.

Self-check: With 100 records and K=10, how many records train and test in each round?

Connects to: Opening Story — Attachment, Mindset, and Ownership, A Study Tool — NotebookLM for Exam Preparation, Metadata Categories for Machine Learning Data.

Metadata — Data About Data

Must-know: Metadata = data about data. Four reasons: discovery (search engine for enterprise data), literacy (common language), trust and compliance (regulated industries), cloud modernization (know what is safe to move). Five dimensions: purpose, date and time, ownership, methodology, location. Five types: descriptive, structural, administrative, technical, statistical.

Top pitfall: Metadata written but never read — the common language only works if producers and consumers both use it.

Self-check: Name the five dimensions of metadata.

Connects to: Metadata Categories for Machine Learning Data, Metadata Management and Lineage Tracking, Flashcard Recap — Data Integration, Ingestion, and Pipelines.

Metadata Categories for Machine Learning Data

Must-know: Six categories: artifact, data set, feature, label, ML model, pipeline. ML model metadata: training parameters, evaluation metrics, prediction examples, data set versions, pipeline outputs, weight file references. Git + DVC (data version control) + MLflow give comprehensive versioning.

Top pitfall: Neural networks (CNN, RNN) do not store labels — only features; traditional ML relies on label metadata. Forgetting licensing in dataset metadata is a legal risk.

Self-check: What does the pickle file carry?

Connects to: Metadata — Data About Data, The ML Metadata Store — Why and How, Metadata Repository, Registry, and Store.

The ML Metadata Store — Why and How

Must-know: Metadata needs dedicated storage (Oracle's user_objects in the system tablespace = the data dictionary). A metadata system minimally tracks feature definitions, versions used in model definitions, and which trained models use them. The ML metadata store is NOT the same as data warehouse or RDBMS metadata.

Top pitfall: Starting a data science project without a solid metadata system and regretting it later; confusing the ML metadata store with warehouse/RDBMS metadata.

Self-check: What should a metadata system minimally keep track of for features?

Connects to: Metadata Categories for Machine Learning Data, Metadata Repository, Registry, and Store, Metadata Management and Lineage Tracking.

Metadata Repository, Registry, and Store

Must-know: Repository = objects + relationships (GitHub, post-commit hooks). Registry = checkpointed important metadata, always specific (model registry lists production models). Store = shopping for metadata (search, compare, pick). The store treats the pipeline or the model as a first-class object depending on what is in the center.

Top pitfall: Putting every draft into the registry — the registry's purpose is to find the important thing fast; using 'store' and 'repository' as synonyms in a case-study answer.

Self-check: Given a case study, when would you recommend a registry over a store?

Connects to: Metadata Categories for Machine Learning Data, The ML Metadata Store — Why and How.

Metadata Management and Lineage Tracking

Must-know: Metadata management = capture, organize, maintain. Lineage captures origins, movement over time, and what happened between creation and the present state. Three types: forward (source to destination), backward (origin), horizontal (across systems). Experiment tracking keeps runs, parameters, scripts, and configuration files.

Top pitfall: Tracking results without tracking configuration — an experiment you cannot reproduce is not an experiment.

Self-check: What is the difference between forward and backward lineage?

Connects to: Metadata — Data About Data, The ML Metadata Store — Why and How, Metadata Categories for Machine Learning Data.

Exam Guidance Summary

Must-know: Expect questions on ML experiments and ML metadata, including repository/registry/store case studies, the four experiment types, K-fold/LOOCV/LPOCV, and the six-step validation checklist.

Top pitfall: Answering a case study without naming the option and walking through the setup.

Self-check: What would you recommend — repository, registry, or store — and how would you go about it?

Connects to: Machine Learning Experimentation, Metadata Repository, Registry, and Store, Metadata — Data About Data, A Study Tool — NotebookLM for Exam Preparation.

Key Industry Applications

Must-know: Git versions code, DVC (data version control) versions data sets, and MLflow tracks experiments and models; relational databases keep metadata in dedicated repositories.

Connects to: Metadata Categories for Machine Learning Data, The ML Metadata Store — Why and How.

Was this lecture useful?

Loading comments…
🤖

BitsNotes AI Assistant

Subject Notes Assistant

Configure AI Key

Select Provider & API Key
🔑 Enter API key above to fetch live models from provider, or enter model name manually.
OpenAI-Compatible API Support

Choose any provider preset (Gemini, DeepSeek, Kimi, GLM, MiniMax, Qwen, OpenAI, Groq, Ollama, etc.) or enter a custom endpoint URL.

Security & Privacy First

Your API key is sent directly from your browser to your specified provider. BitsNotes servers never store or see your key.