Tableau Desktop: Data Connection and Preparation
Prerequisite Knowledge
This lecture builds on the following concepts from earlier lectures. If any feel unfamiliar, review the linked notes before proceeding.
Previously Covered in This Subject
- Tableau student licensing and setup — covered in Lecture 3 (the one-year student license route, Tableau Public versus Desktop)
- Tableau and Salesforce — covered in Lecture 3
- Stories in Tableau: sheets, dashboards, and story points — covered in Lecture 3
- Power BI versus Tableau — covered in Lecture 3
- Hierarchies in Tableau — covered in Lecture 3
- Data types: the four building blocks — covered in Lecture 4 (categorical, numerical, ordinal, time series)
- Numerical data: discrete and continuous — covered in Lecture 4
- Tableau Desktop as a desktop visualization mainstay — covered in Lecture 5
6.1 Module 2 kickoff: from theory to hands-on Tableau
6.1.1 Where the course stands and how this module works
Module one is done — that was the theoretical part of the course. Module two starts today, and from here on the sessions are mostly hands-on: a combination of live demos and example sharing, with Tableau Desktop used throughout. The plan for the next few classes is to spend most of the time working inside Tableau itself.
Hook: Most people think you learn a BI tool by watching demos. This module takes the opposite bet: you learn it by opening the tool and clicking around yourself. By the end of these sessions, the goal is not to recite a feature list — it is to open Tableau on your own and build something from scratch, and to walk into an interview and speak about the tool with confidence.
The instructor made a point of preparing much more than the couple of lines the course handout listed for this session; the reasoning is simple — the goal is not to memorize a feature list but to understand the product deeply enough that you can open it and independently create things. A second, explicit goal: by the end of these sessions you should be able to walk into an interview and speak about this tool with confidence. Everything builds toward dashboarding and storytelling, which is where the course is headed step by step.
Because the session is demo-heavy, the instructor warned that the walkthroughs might jump around a little (finding data files, dragging, dropping) and asked for patience. The session was also designed to be re-watchable: the class was told they could always review the video again. The overall teaching stance for the rest of the course: do not limit the content to the handout; expect the sessions to go well beyond it, because the aim is that you can create things on your own afterward.
6.1.2 The Tableau product family
Tableau is not a single tool but a cluster of products, and it helps to know what each one is for, even if you will not use all of them:
- Tableau Desktop — the individual analysis tool. If you are doing something for yourself at an individual level, this is the one. Every demo in this course uses Desktop. It is a complete tool by itself; you do not need the others to work with it.
- Tableau Server — the collaboration layer. You publish your reports and outcome worksheets here, and other people can see them on the server itself.
- Tableau Reader — strictly for reading and viewing. It cannot alter, create, or change any content that Desktop or Server produced. It is a view-only companion.
- Tableau Prep and Tableau Data Management — add-ons used when working on very large programs and huge datasets. From an individual Desktop point of view you do not need them.
- Tableau Mobile — the mobile companion.
- Tableau Public — the free platform anyone can use. It does not come with all the features of Desktop, but it is a legitimate free starting point.
- Tableau Extensions — the ecosystem of add-on extensions.
A useful way to sort this family in your head: Desktop creates, Server shares, Reader views, Prep cleans at scale, Public lets anyone start for free. The product names appear all the time in job postings and interview conversation, so even a one-line sense of each one is valuable.
Scope: All of this course's work happens in Tableau Desktop. Server, Reader, Prep, and Data Management exist for team and enterprise scale — you do not need them installed to follow along, and you should not feel you are missing something when the demos only use Desktop.
Real-world: Tableau is owned by Salesforce, and the class included someone with a Salesforce background who was invited to share any real-world Tableau associations from their work during the course. The official website walks through every product and offers a demo for each; that is the best place to browse the full list.
6.1.3 Installing Tableau: trial and student license
There are two ways to get Tableau on your machine. The free trial lasts only 14 days — probably not long enough to work through a course, so treat it as a quick look. The better route for students is the dedicated student page: you enter your details (name, the institute you study at, what you will use it for), and the site emails you a license key valid for one year. You need your student email address for this. One classmate had already applied, and one student had already installed Tableau by the time of this session.
The instructor's strong advice: actually install it rather than just watching. Desktop has many more features to play with than Tableau Public, and the one-year student key is free.
6.1.4 How to practice
The follow-on advice for learning: watching demo videos (there are plenty on YouTube) is not enough — you only really learn this kind of tool by hands-on use. Start small with your own day-to-day data: your expenses, your investments, your portfolio. Create an Excel file, connect to it, drag your features and fields around, and see how the visualization reacts. That small start is how everybody learns the tool.
Q: Has anyone already downloaded Tableau? A: One student had. For everyone else the message was: get the one-year student license started soon, because the coming sessions assume you can follow along on your own copy and play with the data yourself.
The practice loop to build this week: pick a tiny personal dataset (your monthly spending works well), connect to it, drag a couple of fields onto the shelves, change the chart type, and ask one question each time ("what happens if I drag this field over here?"). Ten minutes of that beats an hour of passive video watching, because the mistakes you make — and undo — are what stick.
6.1.5 The universal BI workflow
Before touching Tableau, it is worth seeing the shape that every BI tool shares. Whether it is Tableau, Power BI, or Qlik, the workflow has the same three functions:
- Connect to the data — point the tool at your data source.
- Drag and drop fields — place fields on the shelves; the visualization builds itself.
- Save, interact, or publish — keep the result, keep exploring it, or share it.
Every tool does exactly this; what changes between tools is the template, the GUI, and the ease of use. Some tools are niche, some are very light. The instructor had used another BI tool only a few days earlier (the demo shown in the previous class) and it took no time at all to get familiar with it and start dragging and building — evidence that once you know this workflow in one tool, the others come fast.
Exam note: Interview readiness is an explicit goal of this module. Expect questions about product vocabulary (Desktop vs Server vs Reader, Public vs Prep) and about the workflow itself — connect, drag and drop, save or publish. Knowing the three-step universal BI workflow is the single safest answer you can give across Tableau, Power BI, or Qlik.
Recap: Module two shifts the course from theory to hands-on Tableau Desktop. The product family has one tool you actually need (Desktop) plus collaboration, viewing, prep, and free variants. Get the one-year student license, practice on your own small data, and remember the universal BI workflow that every tool shares — connect, drag and drop, save or publish.
6.2 Connecting to data: the three entry points
6.2.1 The starting screen
When Tableau opens, the very first screen is a data-source picker with three categories:
- File — connect to a file you have on disk.
- Server — connect directly to a database server.
- Saved data sources — reuse data sources Tableau ships with or that you have saved before.
This screen is the gateway to everything: no data, no visualization. You pick one of the three entry points, point Tableau at your data, and the fields appear ready to use.
Intuition: Think of the starting screen as the reception desk of a building. You arrive with your data and state where it lives: in a folder on your computer (File), in a shared database somewhere on the network (Server), or already prepared in a data source you (or Tableau) saved earlier (Saved data sources). The reception desk sends you to the right door, and every door leads to the same floor — the data pane with your fields in it.
There is a live demo of the Excel flow: pick the Excel file, and the sheet's columns arrive as fields (Name, Age, Gender), ready to drag. And a useful habit: Tableau's undo behaves like MS Word undo — you can walk back through everything you just did to return to the starting screen.
6.2.2 File-based connections
The file category covers Microsoft Excel, text files (CSV is the common one), statistical files, and PDF files. These are the formats people actually use in practice, Excel and CSV first, with PDF as a capable extra (more on PDFs in the Data Interpreter section). The flow is the same for any of them: pick the file, the columns appear as fields, and you can immediately start dragging fields to build graphs.
6.2.3 Server-based connections
The server category is where the breadth of Tableau shows: the list contains roughly 77 kinds of connectors, ready to use. The instructor clicked through and picked a few at random: Databricks, MySQL, Google Sheets (it lets you log into your Google account and pull the data from there). On top of the ready-made connectors, Tableau offers JDBC and ODBC connections, which cover any database that supports those standards — which is essentially every mainstream database.
The live demo used MySQL on the instructor's laptop: entering the usual credentials (server address, username, password) connected Tableau directly to the database, and every table in it appeared in the data pane, ready for drag-drop and joins between tables. A practical note from the demo: the connection dialog differs per database. MySQL asks for one set of fields, SQL Server asks for different parameters, and so on. But the knowledge you need is always the same three things: the name of the database, the username, and the password. With those, you can point Tableau at any server.
Worked example: the Excel connection demo. The instructor picked the Excel file entry point and chose a small spreadsheet. The moment the file loaded, its sheet appeared in the data pane with three fields: Name, Age, Gender — the spreadsheet's columns, one field each. Each field carried its type icon (ABC for the text columns, hash for the numeric age column). No setup, no mapping step: the columns simply arrived as fields, ready to drag onto a shelf. Sense-check: a spreadsheet column is the unit Tableau calls a field, so the number of fields you see equals the number of columns in the source table.
Worked example: the MySQL server demo. The instructor switched to the Server category, chose MySQL, and entered the credentials: the server address, a username, and a password. Tableau connected straight to the running database, and the data pane filled with every table from that database. Each table could be dragged in; tables could also be joined with each other, because they were all reachable from the same connection. The key transferable point: the dialog for MySQL is not the dialog for SQL Server, but the information you must supply is the same three items — database name, username, password. Sense-check: a server connection is nothing more than the database name, the user, and the password, so these three facts are all you need to memorize for any connector.
The two demos cover the extremes of the entry-point spectrum — a plain file on disk and a live database server — yet both end at the same place: a data pane full of fields waiting to be dragged.
6.2.4 Saved data sources
The third category covers predefined data sources that come with Tableau itself — no setup needed. The Superstore sample is the one used all through this session: it brings tables like Order, People, and Return with many fields to play with (dimensions, dates, measures). Saved data sources are the fastest way to start experimenting when you do not have your own data at hand.
6.2.5 When to use file versus server
The choice between a file and a server is a scenario question, and the instructor gave clear decision rules:
Use file-based when:
- The data will not change frequently — a monthly extract, a quarterly report, a one-time analysis.
- You have downloaded the data and you do not need real-time values.
- The data is not sensitive — you will be holding it offline and uploading it where needed.
Use server-based when:
- The data is changing constantly, with users entering records.
- You need real-time graphics that track the current state.
- The data is sensitive — it sits on a password-protected server with protective measures in place.
So the mental shortcut: static, not sensitive, no real-time need → file; changing, sensitive, real-time need → server.
| Scenario | File | Server |
|---|---|---|
| Data changes | Rarely (monthly/quarterly snapshots) | Constantly (live records) |
| Real-time need | No | Yes |
| Sensitivity | Low (held offline) | High (protected server) |
| Typical case | Personal analysis, one-off reports | Shared business databases, dashboards |
Recap: Three entry points connect data into Tableau — file, server, and saved data sources. Files (Excel, CSV, PDF, statistical) are for static, non-sensitive data you hold offline; servers (about 77 connectors, plus JDBC/ODBC) are for changing, sensitive data needing real-time access; saved data sources like Superstore are the fastest way to start experimenting. The knowledge needed for any server connection is always the same: database name, username, password.
6.3 Live connections versus data extracts
6.3.1 What a live connection is
Every time you connect to a source, Tableau asks which kind of connection you want: live or extract. A live connection is a direct connection to the underlying data. If something in the underlying data changes, your reports change too — the moment you refresh, the new value reflects in the graphics in real time. The underlying source is the single source of truth; the report simply reads from it.
Intuition: A live connection is like reading the news from a live website — the page you look at is whatever the site shows at that instant, and when the site updates, your refreshed view shows the new story. An extract, in contrast, is like printing today's newspaper and reading the printout next week: comfortable offline, but it will never carry news that broke after the print run.
6.3.2 What an extract is and why people use it
A data extract is a copy of the data taken as of a particular time and date. It is a subset — a snapshot you carry with you. Once you are working on an extract, changes in the source data do not affect your graphics at all; you are working offline against the saved copy.
Why would anyone choose a snapshot over live data? The advantages the instructor listed:
- You can download a big dataset and work offline — on a plane, at an airport, anywhere without connectivity.
- It is faster — no live connection overhead, no round trips to the database.
- Performance is better for heavy graphics work.
The flip side: live connections are heavier on performance because the tool constantly deals with real-time data. So the trade-off is freshness and space-safety versus speed and portability.
6.3.3 Worked demo: the same Excel through both modes
The instructor showed the difference end to end with a tiny dummy Excel file.
Setup. The Excel file had three columns — Name, Age, Gender — with nine rows of dummy names (A, B, B, and so on). Tableau was connected live, and the columns arrived exactly as they were in Excel. Then two fields were dragged into the view — Name and Age — producing a simple bar chart of the ages.
Live behavior. In the Excel file, one cell (the row labelled ZZZ) was changed to the text "changing for demo", and the file was saved. Back in Tableau, the chart did not update on its own — but the moment the report was refreshed, the change appeared in the graphic. That is the live behavior: refresh propagates the source change into the report.
Extract behavior. The connection was then switched to extract and the extract was saved. Again the Excel file was edited — this time to "changing second time for the demo" — and saved. Back in the report, refresh changed nothing. The graphics kept the old values, because the extract is the copy you chose to work from, not the Excel file itself.
Back to live. Switching the connection back to live and refreshing finally brought "changing second time for the demo" into the chart. One mode propagates changes, the other does not — that single difference is the whole story.
Worked example: the ZZZ cell, end to end. The test file was a tiny table of three columns (Name, Age, Gender) and nine rows of dummy names, where one row was labelled ZZZ. In live mode, the row's cell was edited to the text "changing for demo" and the file saved. The bar chart of Age did not move on its own; pressing refresh brought the new value into the graphic — proof that live reads the source. The connection was then switched to extract and saved. The same cell was edited again, this time to "changing second time for the demo", and the file was saved. Refresh now changed nothing; the chart kept the first edit's value, because the extract had captured the data at the moment it was saved. Finally, switching back to live and refreshing showed the second edit too. The experiment isolates one variable: refresh propagates source changes on a live connection, and does nothing on an extract. Sense-check: whichever mode you are on, the refresh button is the honest test — if the value updates, you are live; if it stays frozen, you are reading a saved copy.
The instructor's summary line: extract is a copy as of that time and date; live is something which can change — and the refresh test is the way to check which one you are on.
6.3.4 Spotting the mode: database icons
There is a small visual tell you might miss at first glance: a live connection shows one database icon on the data source, while an extract shows two database icons. If you ever forget which mode you are in, look at the icon count before you refresh.
Scope: The live-versus-extract choice is about freshness versus performance, not right versus wrong. An extract is the right call for large, slow-changing data you work with offline or on heavy charts; live is the right call when the numbers must track the current state of a changing source. The pitfalls are misreading which mode you are in (check the icon count) and expecting an extract to update when the source changes — it never does until you create a fresh one.
Recap: Live connections read straight from the source, so refresh brings changes in; extracts are saved copies from a moment in time, so refresh changes nothing. Choose live for freshness and space-safety, extract for speed and offline portability. One database icon means live, two means extract — and the refresh test tells you for certain.
6.4 Data types in Tableau
6.4.1 The six field types and their icons
Whenever you export or connect data, each field carries one of roughly six data types, and Tableau shows the type with an icon at the top of the field, so you can read the schema at a glance:
- String — text, shown with an ABC icon. If the field contains alphabets, it is a string.
- Date — shown with a calendar icon (Order Date and Ship Date in the sample data are dates).
- Date-time — a date with a time component.
- Number — shown with a hash (#) icon. This covers both Number (decimal) and Number (whole).
- Boolean — true/false values, shown as Y and N.
- Geographic — shown with a globe icon; location fields like country and city fall here.
Intuition: The type icon is the field's identity badge. Before you know what a field contains, the badge tells you how Tableau will treat it: a hash badge means "this can be added, averaged, summed", an ABC badge means "this is a label", a calendar badge means "this can be sliced by year, quarter, or month". A complex table with many fields is readable in seconds once you can read the badges.
Not every table has every type, but these six are the broad categories you will meet in any database. The icons matter because when a complex table arrives with many fields, the icon is how you immediately recognize what kind of field you are holding.
6.4.2 Changing a field's data type
The type Tableau infers is not final. If you upload data and decide a field would be better as something else — say you would rather keep a numeric-looking field as a string — you can change it: right-click the field, go to change data type, and pick the new type. The instructor showed converting an alphabetic field to boolean, noting that this particular change is not logical for that data, but the mechanism is exactly this. Your call on what the right type is.
6.4.3 Number subtypes and special roles
Numbers come in decimal and whole variants. And geographic data is special: location fields get special geographical roles (country, state, city, postal code and so on) that later power map visualizations. Those roles are covered later in the course; for now, recognize them by the globe icon.
Pitfalls:
- The inferred type is a guess, not a guarantee: an ID column of digits may arrive as a number when you want it kept as text (leading zeros like 007 disappear the moment it is treated numerically — change it back to string).
- Changing a type to something illogical (like alphabetic names to boolean) is possible but meaningless — Tableau allows the change, but the resulting field makes no sense to analyze.
- A date stored as text (ABC icon) cannot be drilled by year/quarter/month until you change its type to date.
Recap: Every field carries one of six data types — string (ABC), date (calendar), date-time, number (hash), boolean (Y/N), geographic (globe). The icon is the at-a-glance identity of the field, types are changeable per field via right-click, and numbers plus geographic roles have subtypes that matter later for maps.
6.5 Data Interpreter: automatic data cleaning
6.5.1 What it is and when it appears
Tableau has some built-in intelligence for deciding on and cleaning up data. Give it data that is not very structured and it applies its own logic to tidy it. This is a feature the instructor flagged as under-shared in most demos — and honestly Power BI is more mature on auto-filtering and auto-correction, but Tableau still has real capability here.
Two things about when the Data Interpreter shows up:
- It appears only when Tableau detects a discrepancy — untidy data where it is struggling: merged cells, nulls, blank fields, gaps. For clean, straightforward, structured data the option does not appear at all, because there is nothing to interpret.
- It may not activate for very small tables (few rows and columns), in which case you clean the data offline and bring in a ready-made clean file.
It is always optional — never imposed on you. You can decline its interpretation and clean the data yourself.
Intuition: The Data Interpreter is like an editor who reads a messy handwritten page: it spots that titles are titles, headers are headers, and numbers are numbers, and types up a tidy version. It only offers to help when the page is actually messy — a clean typed page gets no offer. And like a good editor, it always hands you a marked-up copy of what it changed, so the final call stays with you.
6.5.2 Worked example: merged header cells
The demo file was an Excel sheet whose headers used merged cells: the 2017 header spanned two columns, 2016 spanned two, and 2015 spanned two. When Tableau first read the file, it only read the first row of each merged cell — so each year's second column came in as null. It literally could not see the repeated year.
Toggling Use Data Interpreter fixed it instantly: the interpreter built the logic on its own, recognized the merged headers, and propagated the year values — filling 2017, 2016, and 2015 across the three affected rows. The instructor undid and re-enabled it to show the before/after: undo brought the nulls back, and enabling the feature filled them again. The lesson: when you see nulls that look like they should not be null, the Data Interpreter is the first thing to try.
Worked example: merged headers filled automatically. The sheet had a header row of three merged cells — the text 2017 spread across two columns, 2016 across two, and 2015 across two. Tableau's first read only looked at the leftmost cell of each merged pair, so the right-hand columns of each year pair arrived as null: the year label simply was not repeated in those columns. With Use Data Interpreter toggled on, Tableau recognized that each merged label applied to both columns and propagated the year values into all three affected rows — 2017, 2016, and 2015 each appearing where the blank sat. Toggling the feature off brought the nulls back, and on again refilled them, which proved the interpreter, not manual editing, did the work. Sense-check: a merged cell visually spans columns but physically stores its text once; the interpreter's job was to copy that one stored value into every column the merge covered.
6.5.3 The review report
Data Interpreter does not act silently. A Review Results report tells you exactly what it did to your data, so you stay in control:
- Red cells = what it decided are column headers.
- Green cells = transactional data.
- Dark green highlights = the specific places where it made changes.
In the demo the review pane flagged 27 such changes. The interpreter can detect and bypass common problems on its own: titles, notes, footnotes, and empty cells. If the report shows two separate tables in one worksheet (separated by row or column gaps), it splits them and hands you two tables to play with instead of one mangled one. And the final decision is yours: if you like the interpretation, proceed; if not — "your interpretation is wrong, I will do my own data myself" — you decline and move on.
6.5.4 Worked example: reading a PDF
PDFs are a special case that the Data Interpreter makes practical. The demo used a PDF report of Amazon stock containing a historical data table.
When you pick a PDF, Tableau asks how much of it to read: one page, all pages, or a selected page. The demo chose all pages, and Tableau came back with a separate table for each page — page 1, table 1; page 2, table 1, and so on. Page 1's table was the interesting one: before interpretation, the month-year column was blank. Enabling the Data Interpreter automatically filled the January, February (and so on) month-year values, removed a first line that was of no use, and turned anything it could not use into nulls. Its own report then explained: this line was treated as a header, this table is more or less clean. Page 2's table came in clean already.
The honest framing from the demo: the interpreter's guess is not necessarily accurate — "not necessarily to be accurate, but one can actually fix it". It is a starting point that saves you from doing the entire cleanup by hand, and you can always correct the result afterward.
Worked example: the Amazon stock PDF. The source was a PDF report on Amazon stock with a historical data table. Tableau offered three reading scopes — one page, all pages, or a selected page — and the demo chose all pages. The result: one table per page (page 1 → table 1, page 2 → table 1). Page 1's table had a blank month-year column; enabling the Data Interpreter filled January, February, and the following months automatically, dropped a first line that carried no useful data, and turned unreadable fragments into nulls. The interpreter's own report labeled which line was treated as a header and confirmed the rest of the table was clean. Page 2 needed nothing. Sense-check: a PDF stores text and numbers visually, not as a data table, so extraction is a guessing game — which is why the interpreter is a starting point to correct, not a final answer.
Scope: The Data Interpreter is a helper for untidy sources, not a magic wand. It only appears when it detects discrepancies; it may skip very small tables; and its guesses can be wrong. For genuinely structured data it never shows up, and for small messy files the cleaner path is to tidy the data in Excel first and connect to a ready-made clean file.
Recap: The Data Interpreter auto-cleans untidy data — merged headers, blanks, titles, footnotes — but only appears when it detects a discrepancy, is always optional, and always reports what it changed (red headers, green data, dark-green edits). PDF reading relies on it heavily, and the result is a starting point you can fix, not a guaranteed-accurate read.
6.6 Viewing data before building
6.6.1 The table view
Before building any visualization, look at the data. Every table in the data source can be opened in a table view that shows every field and every value — a real table, not a preview. For the Superstore sample, opening the Order table shows its full set of fields (Row ID, Order ID, Order Date, Ship Date, Ship Mode, Customer ID, Segment, Category, Sub-Category, Product Name, Region, City, Country, Postal Code, Sales, Profit, Discount, Quantity, Profit Ratio, and more) with all values. This is the quick way to know what you are actually working with before you start dragging.
Worked example: the Order table view. Opening the Order table rendered a real table with one column per field — Row ID, Order ID, Order Date, Ship Date, Ship Mode, Customer ID, Segment, Category, Sub-Category, Product Name, Region, City, Country, Postal Code, Sales, Profit, Discount, Quantity, Profit Ratio, and more — and one row per order record, every value visible. Before a single field was dragged onto a shelf, this view answered the three questions that matter: what columns exist, what each column holds, and how many rows there are. Sense-check: the table view is inspection only — it shows the data source as-is, and nothing you look at here changes the data.
6.6.2 Filtering the view
A full table can be overwhelming. The view supports filtering: you can narrow it down — for example, to a slice of the fields — so you are looking at only what matters. Filtering the view is purely for your inspection; it does not remove fields from the data source.
6.6.3 Hide and unhide fields
Once you have looked at the data, you can decide which fields you want to keep in play. Right-click a field and choose Hide — the field disappears from the data pane. This is useful when the field list is heavy and you do not want everything in your visualization.
The trick is getting it back: use Show Hidden Fields in the data-pane menu, and every hidden field reappears grayed out. Then right-click and Unhide the ones you actually want. So hiding is reversible, and you can pick and choose your working field set freely.
Pitfalls:
- Confusing view filtering with deleting data — filtering the view only changes what you inspect; every field and row stays in the source.
- Forgetting how to bring a hidden field back — remember the Show Hidden Fields option in the data-pane menu, which restores hidden fields grayed out, ready to unhide.
- Hiding a field because the list feels heavy, then wondering why a chart is missing it — hiding is per your working set, not per the workbook's capability.
Recap: Look before you build — the table view shows every field and value of a table, view filtering narrows what you inspect without touching the source, and hiding fields (reversible via Show Hidden Fields) trims the data pane to your working set.
6.7 Column formatting: rename, reset, copy, split
6.7.1 Rename and reset
Field names can be changed two ways: double-click the field name, or use the right-click menu. The demo renamed the ID field to "My ID". Renaming is for your working convenience, but the instructor's aside is worth keeping: a team may well reject ad-hoc naming conventions, and you might forget the original name — so Tableau provides Reset Name, which restores the original name the field had at the source. Rename freely, reset freely; the source name is always one click away.
Intuition: Renaming a field is like putting a sticky label on a storage box. The label is for your convenience — "the one with my totals" — but the box itself still says what it always said. If the sticky label confuses people or you forget what you renamed, peel it off: Reset Name brings back the original label, one click, no harm done.
6.7.2 Copy a column out to Excel
You can copy a column just like in Excel: select the column, press Control+C, and paste into another application with Control+V. The demo pasted a column of IDs into a fresh Excel sheet (using the Copy Value option to paste plain values), and Control+A selects everything in the view. This is a quick way to take any column, bring it somewhere else, and keep it for reference.
6.7.3 Automatic split
Some columns carry two ideas in one cell, and Tableau can split them. Right-click the field and choose Split: Tableau detects where the parts separate and creates two new columns automatically.
The demo used Ship Mode, whose values look like "Second Class", "First Class", "Standard Class". After split, the data source gained Ship Mode (Split 1) containing the first word (Second, First, Standard) and Ship Mode (Split 2) containing "Class". If you do not want the split, delete the new columns — everything is reversible.
Worked example: automatic split of Ship Mode. Ship Mode held values like "Second Class", "First Class", and "Standard Class" — each cell carries a ranking word and the word Class. Right-clicking the field and choosing Split made Tableau detect the break point (the space) and create two new fields automatically: Ship Mode (Split 1) with the first word (Second, First, Standard) and Ship Mode (Split 2) with "Class". The original Ship Mode field stayed intact, and the two new columns could be used in visualizations on their own. Deleting the new columns removed the split and left the source exactly as before. Sense-check: an automatic split reads the cell contents, finds a repeated separator, and chops each cell at the same place — one separator, one boundary, two columns.
6.7.4 Custom split
Automatic split guesses; Custom Split lets you say exactly how. Right-click the field, choose Custom Split, and specify the separator character and how many pieces to make.
The demo used Order ID, whose format is something like CA-2017-12345 — a state code, a year, and a number separated by hyphens. With a hyphen as the separator, Tableau asked whether to split only the first field or all fields. Splitting the first field produced just the "CA" part. Splitting everything produced three columns — Split 1 = CA, Split 2 = the year, Split 3 = the number — each usable in a visualization of its own.
Worked example: custom split of Order ID. Order IDs look like CA-2017-12345: a state code (CA), a year (2017), and a sequence number (12345), joined by hyphens. Automatic split would guess at the boundary, so the demo used Custom Split with the hyphen as the separator. Tableau then asked a key question: split only the first field, or split all of them? Splitting just the first produced one new column with the "CA" part. Splitting everything produced three columns: Split 1 = CA, Split 2 = the year, Split 3 = the number — every part of the ID now a standalone field usable in a chart or a filter of its own. Sense-check: custom split with a chosen separator and piece count gives you exactly the columns you asked for; automatic split gives you what Tableau guessed.
The automatic and custom splits are siblings with one difference: automatic split guesses the separator for you, custom split takes your exact instructions. The rest of the workflow — create the columns, use them, or delete them if unwanted — is identical.
This whole area is what the instructor called data prep: before you move into graphics, you verify the data is understood — look at the tables, join whichever tables you want, decide whether you need all fields, check the data types, split what needs splitting — and only when the data is ready do you step into the visualization mode.
Pitfalls:
- Renaming fields with ad-hoc names in a shared team workbook — a team may reject your naming convention, and you may forget the original name; Reset Name is the escape hatch.
- Expecting an automatic split to know your intended boundary — it guesses from a repeated separator; use Custom Split with an explicit separator and piece count when the format is specific (like Order ID).
- Splitting a column that is better left whole — both splits are reversible; delete the new columns and the source field is untouched.
Recap: Column formatting is data prep: rename fields (Reset Name restores the source name), copy columns out with Control+C/Control+V, and split combined columns — automatic split guesses the boundary, custom split takes your separator and piece count. Everything here is reversible and nothing changes the source data itself.
6.8 Sorting fields, not data
6.8.1 Sorting the field list
Whatever Excel file or database you connect, the fields arrive in the source's order — Row ID, Order ID, Order Date, and so on, exactly as the table had them. Sometimes you want the field list sorted for easier navigation. From the data pane menu you can sort the fields A to Z — the demo produced Category, City, Country, Customer ID — or Z to A (descending), which puts the last fields first.
There is an option-set nuance here. The first two options (A to Z, Z to A) are the ones to choose when you are working with one table. When you have multiple tables joined (Order plus People plus Return, for example), the per-table sorting options apply the sorting inside each table, in sequence, in one shot — so each table's fields are sorted within that table. Either way you can always go back to the data source to restore the original order.
6.8.2 Sorting data at the column level
Sorting the field list is not the same as sorting the data. For the data itself, click a column header in the table view: first click sorts ascending, second click descending, and a third click returns to the original order. You can sort the data by any field this way — the demo sorted by Ship Date.
6.8.3 The fields-versus-data distinction
Q: Isn't sorting the fields the same as sorting the data? A: No — be clear on this one. Here we are only sorting the columns, the field names, so the data pane is easier to navigate. Sorting the data itself happens later, from the visualization: click a column header to sort ascending, click again for descending, and a third time returns to the original order.
This is exactly the kind of distinction the instructor flagged explicitly — "I am not saying sorting data. Just be clear there. I am talking about sorting the columns." The two features sit in different places and do different jobs.
Worked example: the A-to-Z field sort. In the data pane menu, choosing A to Z rearranged the field list alphabetically, producing the order Category, City, Country, Customer ID — the navigation order of the pane changed, nothing else. The rows of data were untouched: the table's records still sat in their source order, and any chart built from the fields still read the same values. Choosing Z to A reversed the list, and returning to the data source restored the original field order. Sense-check: the sort applied to the shelf of field names you pick from, not to the values those fields hold.
Pitfalls:
- Thinking sorting the field list sorted the records — the instructor explicitly flagged this confusion: the data pane sort only reorders the column names for navigation.
- Sorting data by clicking a column header and forgetting the third click rule — first click ascending, second descending, third returns to original.
- Sorting a field list expecting it to affect charts or filters — field-list order is purely cosmetic; data order only changes when you sort the data itself.
Recap: Two different sorts exist: the field list (A to Z / Z to A from the data pane — cosmetic navigation, per-table options when tables are joined) and the data itself (click a column header: first click ascending, second descending, third original). Sorting fields is not sorting data.
6.9 The visual interface: a guided tour
6.9.1 Menu bar and toolbar
The menu bar is the top strip: File, Data, Worksheet, and the rest. Below it sits the toolbar, where the instructor was doing undos and formatting throughout the session. The toolbar carries the everyday actions:
- Undo / redo — works like MS Word; walk back through anything.
- Save, new dashboard — create a new dashboard directly from the toolbar.
- Clear sheet — empties the whole worksheet; the view goes blank. Use it when you have made a mess and want to start over.
- Duplicate sheet — copy an existing sheet when you want to tweak one thing but keep most of it.
- Ascending / descending sort — sorting from the top menu.
- Fit options — how the visualization is scaled on the page: entire view (compress everything, big or small, to fit the whole page), standard (default sizing), fit width, or fit height. The demo toggled between entire view and fit height to compress the view.
- Fix axis, highlight control — more advanced icons the course will get to; they are among the most commonly used.
Intuition: Think of the toolbar as the dashboard of a car: the everyday controls you reach for constantly — undo, save, clear, duplicate, sort, fit — all in one row, always visible. The menu bar above it is the full manual: everything exists there, but most people drive with the dashboard.
6.9.2 Shelves and cards
Two pieces of Tableau vocabulary you will hear constantly:
- Shelf — the rows and columns areas are called shelves because that is where you keep things: you drag a field there, place it, and the visualization uses it. The instructor's phrasing: "shelf is where we keep things."
- Card — the panels used to format the visuals. The marks card (color, size, shape, label), filters, and pages are cards. This is where you style the marks on your chart.
Next to them is the Show Me panel — the grid of graph types Tableau can offer. Click a thumbnail and Tableau rebuilds the current view as that chart type.
6.9.3 Sheets, dashboards, and stories
Tableau organizes work in three escalating containers:
- Sheet (worksheet) — one individual visualization built from the data. You add new worksheets one at a time.
- Dashboard — a combination of individual sheets. If you have employee data, performance data, and sales data, you build one sheet per view, then place the sheets together on a dashboard. A dashboard is nothing but different parameters clubbed together.
- Story — the presentation layer, built on top of dashboards and sheets. A story is a slide-by-slide navigation: sheet one shows the starting point (say, employee count in 2002), sheet two shows the year-wise growth trajectory, sheet three shows attrition — and you walk the management through the flow like a deck, but with live graphics.
The build order matters: individual sheets first, then dashboards, then stories on top of those. Real-world: the class was asked whether Power BI has a story capability; no one could confirm one. Power BI gives you dashboards, but a playable narrative flow with live graphics appears to be a Tableau strength — though feature sets move, so check current versions before relying on that claim.
6.9.4 Data pane, marks card, status bar, and sheet bar
The tour continues with the remaining regions of the window:
- Data pane (left sidebar) — where the tables and fields live: dimensions on one side, measures on another. Any sets or parameters you create later appear here too. The field icons tell the type at a glance (ABC for text, hash for number, calendar for date).
- Marks card — the card used to build coloring, sizing, and labeling.
- Status bar — the strip showing what the view contains, for example "3 rows by 1 column, sum of profit".
- Sheet bar — the tabs at the bottom holding all your worksheets.
- Main worksheet — the big canvas where you drag-drop fields and the graphic appears.
6.9.5 Pills
A field that you drag onto a shelf is called a pill — the name comes from the shape: a field sitting on a shelf looks like a medicine capsule, which is exactly what a pill is. The term is worth knowing because it is used everywhere in Tableau practice. Two rules of pills:
- A blue pill = a discrete field.
- A green pill = a continuous field.
You will hear people say "put the pill on the shelf" — it simply means placing a field into a shelf to build the visualization.
Pitfalls:
- Treating pill color as decoration — blue (discrete) and green (continuous) pills change how Tableau treats the field, so the color carries meaning (details in the next section).
- Mixing up shelf and card — shelves hold the fields that build the chart (rows, columns); cards format the marks (color, size, shape, label).
- Forgetting the build order — sheets come first, dashboards second, stories on top; trying to build a story before its sheets fails.
6.9.6 Naming sheets and workbooks
Every sheet gets a name — whatever you type shows on its bottom tab. In the demo, a view showing sales and profit by category was named "my profit", and that name appeared on the tab. The workbook as a whole gets its name when you do Save As; that name sits on the top bar. Saving is a full save: reopen the workbook and all values, visuals, and hierarchies are restored, with the data source refreshing in the background.
6.9.7 Student questions and answers
Q: What are the three different tabs showing next to Sheet 1? A: They are all worksheets — Sheet (or worksheet), Dashboard, and Story. We will get to them properly once the database part is done: you build individual sheets, combine them into a dashboard, and build stories on top of that. Right now we are still on the data side, which is the foundation.
A second question compared Tableau with a competing product, so the comparison belongs here too:
Q: Does Power BI have a story capability like Tableau? A: Probably not. The class guessed at it and nobody could confirm one. Power BI gives dashboards, but the slide-by-slide narrative flow — something that plays like a presentation with live graphics — seems to be Tableau-specific. Check the latest Power BI features before relying on that.
Exam note: The terminology of this tour — shelf, card, pill, sheet, dashboard, story — is the vocabulary that "will help you a long way" in Tableau practice, and the instructor treated it as interview material. Know the shelf-versus-card split (keep things vs format visuals) and the sheet → dashboard → story build order.
Recap: The interface is menu bar plus toolbar up top, shelves and cards where fields live and get styled, Show Me for switching chart types, the data pane on the left, and three containers in order: sheets (one visualization), dashboards (sheets combined), stories (a playable deck of sheets and dashboards). A dragged field on a shelf is a pill — blue means discrete, green means continuous.
6.10 Dimensions and measures
6.10.1 What a dimension is
A dimension is a field that categorizes or segments the data. Typically these are text, dates, and locations — things you classify with, like customer names, product categories, regions, and country names. Dimensions are qualitative and descriptive: they put labels on things and break the data into groups.
Three behaviors follow from that:
- Dimensions decide the grouping of a chart. "Show me sales by category" — the category dimension creates one chunk per category.
- Dimensions do not impact calculations — they are not aggregated. Grouping and sorting happen based on them, but you never sum a dimension. If you try to sum a country name, Tableau refuses — the record count is what exists for a dimension.
- Bar graphs and charts are typically built around dimensions; they provide the bars' categories.
Intuition: A dimension answers the question "for each what?" — for each category, for each region, for each month. It is the bucket the numbers are dropped into. A measure answers the question "how much?" — how much sales, how much profit. Put the two together and you get a chart: dimensions make the slices, measures fill them with numbers.
6.10.2 What a measure is
A measure is a field that quantifies something — sales, profit, discount, profit ratio, quantity, percentages. Measures are numerical and are the real values in a chart.
Their key behavior is automatic aggregation. When you bring a table containing names and sales, Tableau knows sales is a measure and shows the sum by default. You are not stuck with sum: the instructor showed the aggregation can be changed to average, median, count, and so on — Tableau offers the options and you choose. Measures are the numbers that enter calculations.
6.10.3 The key differences at a glance
| Dimensions | Measures | |
|---|---|---|
| Nature | Descriptive, qualitative | Quantitative, numerical |
| Job | Categorize and group | Quantify and calculate |
| Examples | Customer names, product categories, regions | Sales, profit, quantity, percentages |
| Aggregation | Never aggregated | Aggregated automatically (sum by default, changeable) |
| Role in charts | Define the bars/groups | Define the values |
6.10.4 Example: sales by category
The live example made the pair concrete. Placing the Category field (a dimension — its name happened to be "category" too) on the shelf and Sales (a measure) next to it produced one bar per category; switching to sub-categories produced more bars, one per sub-category. In every case the dimension decides the grouping and the measure supplies the number on each bar. Country names and customer names are dimensions; the figures displayed beside them are measures.
Worked example: one bar per category. The Category dimension was dragged to the shelf, and Sales next to it. Tableau split the records into groups — one chunk per category — and summed Sales inside each chunk, drawing one bar per category with the category's total on it. Switching from Category to Sub-Category kept the same logic and produced more bars, one per sub-category, each with its own summed sales figure. The pattern repeats everywhere: the dimension names the group, the measure gives the number. Sense-check: bar count follows the dimension's values (however many categories exist, that many bars), and bar height follows the measure's aggregation (the sum of sales inside each group).
6.10.5 Why students mix these up
This is a flagged confusion point: students keep interchanging dimensions and measures. The concept is not complex — descriptive versus quantitative, categorize versus quantify — but the terms come up in almost every Tableau conversation, so the instructor treated the distinction as important and repeated the comparison several times. If you take one vocabulary pair away from this session, take this one. The contrast resurfaces throughout the course in every chart you build.
Scope: The dimension/measure split is a behavior rule, not a data-type rule. Tableau assigns roles by default (text and dates usually become dimensions, numbers become measures), but the real test is behavioral: if it can be meaningfully aggregated into a number, it is a measure; if it only labels and groups, it is a dimension. A numeric ID you never want summed is better treated as a dimension — which is why you can change the role by right-clicking the field.
Exam note: The instructor explicitly flagged that students interchange these two terms, called the distinction important, and repeated it several times. Expect the dimension/measure contrast to be tested — descriptive vs quantitative, categorize vs quantify, never-aggregated vs auto-aggregated (sum by default, changeable to average, median, count).
Recap: Dimensions categorize and group (customer names, categories, regions — never aggregated); measures quantify and calculate (sales, profit, quantity — aggregated automatically, sum by default). Dimensions define the bars; measures define the numbers on them.
6.11 Discrete and continuous fields
6.11.1 The two natures of fields
Beyond dimensions and measures, every field also has a nature: discrete or continuous — the same distinction you meet in statistics.
- Continuous means no breakage: an unbroken scale you can slide along. Continuous fields show up green in Tableau.
- Discrete means separate and distinct: individual categories or bins. Discrete fields show up blue.
The color rule is automatic and visible: place fields on a shelf and the pills turn blue for discrete and green for continuous. The instructor's example: a line of a share value per day, changing over time, is always continuous. You can sometimes switch a field's nature, but it depends on the field type — you cannot make a category list continuous in a meaningful way.
Intuition: Think of a ruler versus a staircase. A ruler is continuous — every tiny position between 1 and 2 exists, nothing is skipped. A staircase is discrete — you stand on step 1, step 2, or step 3, and never between. A stock price on a date axis is a ruler: it moves through every value. A bar chart of categories is a staircase: each bar is its own step, and no bar sits between two categories. That is why pills are green on continuous scales and blue on discrete ones.
6.11.2 Examples and where you will see it
You will meet the pairing constantly: date lines and stock-style series run as continuous greens; category bars run as discrete blues. When the drill-down date hierarchy is used (next section), the discrete view gives you year/quarter/month/day stepping, while the continuous view treats the same dates as an unbroken axis. The course will return to discrete versus continuous in the coming weeks with many examples — this session only established the vocabulary and the color coding.
Scope: Discrete/continuous is a nature of the field, and the two rules are: the color is automatic (blue pills for discrete, green for continuous), and switching is limited by field type — dates and numbers can often take either nature, but a category list cannot meaningfully become continuous. When in doubt, ask what the visual is doing: bars and headers want discrete; lines and axes want continuous.
Recap: Every field is discrete (separate, distinct — blue pill) or continuous (unbroken scale — green pill), matching the statistics distinction. Date series and stock lines run continuous; category bars run discrete. The color coding is automatic, and the choice returns throughout the course.
6.12 Hierarchies: drill up and down
6.12.1 What a hierarchy is
A hierarchy is a logical grouping of fields that naturally go together. Instead of dragging country, region, state, city, and postal code onto a view individually every time, you create one hierarchy, drag it once, and then drill up and down through the levels. Tableau shows hierarchies with a special icon in the data pane, so you can spot them at a glance.
The core advantage: you do not have to drag individual fields each time. One hierarchy is one drag, and roll-up/roll-down is a click. The fields still work individually — dragging a field out of a hierarchy does not destroy it — but the hierarchy is there to save work.
Intuition: A hierarchy is like a pre-built nesting doll set: the outer doll is the coarsest level, the inner dolls get finer and finer. You pick up the whole set with one hand (one drag) and decide on the spot how deep to open it (drill down) or how far to close it (drill up). Without the set, you would hunt for each doll separately every single time.
6.12.2 The inbuilt date hierarchy
Dates come with a built-in hierarchy — you do nothing to create it. Drag a date field (say Order Date) and a measure (Profit) onto the view, and Tableau automatically starts at YEAR; the plus sign drills down through QUARTER, then MONTH, then DAY. The demo walked all the way down: year-wise (2018), then quarter-wise sales, then month-wise, then day-wise — and collapsed back up the same way. Whatever granularity the stakeholder wants is one click away.
Worked example: drilling the date hierarchy. With Order Date and Profit on the view, Tableau opened at the YEAR level — one bar per year (for example, 2018). Clicking the plus sign moved the view to QUARTER, then to MONTH, then to DAY, each click re-grouping the same profit values into finer slices; clicking the minus signs collapsed the view back up through the same levels to the year. No new fields were dragged at any step — the hierarchy supplied every level. Sense-check: the number of marks changes with the level (four quarters become twelve months, then many days), but the total profit never changes — drilling only changes how the same total is sliced.
6.12.3 Building your own hierarchy
User-built hierarchies are simple to create: drag one field onto another in the data pane, and Tableau asks for a hierarchy name. The demo grouped Category and Sub-Category into a hierarchy named "my category" — the two fields are always used back to back, so they belong together.
The demo then showed the payoff. A stakeholder asks for numbers by category — one drag of "my category". The same stakeholder then asks for sub-category — one drill-down click, no re-dragging. Without the hierarchy, each request meant dragging fields individually. The hierarchy is user-created, and right-clicking it and choosing Remove Hierarchy turns the grouped fields back into individual ones whenever you want.
Worked example: the "my category" hierarchy. Category and Sub-Category always appeared together in analyses, so the demo dragged Category onto Sub-Category in the data pane; Tableau asked for a name, and the pair became a hierarchy called "my category". A stakeholder asked for numbers by category — one drag of "my category" produced the category view. The same stakeholder then asked for sub-category detail — one drill-down click, and the view regrouped into sub-categories without dragging anything new. The same move works in reverse: right-click the hierarchy, choose Remove Hierarchy, and Category and Sub-Category return to the individual field list. Sense-check: a hierarchy is a saved drag, so the time saved equals the number of times the same field set is reused.
6.12.4 The location hierarchy
The sample data came with a ready-made location hierarchy: country, region, state, city, postal code — fields so closely tied that someone grouped them once. The demo drilled country → region → state → city in one shot (the sample only contains the United States as a country, which is why the demo moved down levels quickly). Any hierarchy can be rolled up or zoomed down the same way.
Worked example: the location hierarchy. The sample data shipped with a hierarchy of country → region → state → city → postal code. One drag of the hierarchy and one click walked the demo from country down to region, state, and city; because the sample contains only the United States as a country, the country level collapsed to a single value and the drill moved quickly through the finer levels. Rolling back up returned through the same chain. Sense-check: a level with a single value still exists in the hierarchy — the drill just passes through it fast.
6.12.5 Adding, reordering, and removing fields
Hierarchies stay flexible:
- Add a field — drag it onto the hierarchy; the demo added Manufacturer to "my category".
- Reorder — drag levels up and down; moving Manufacturer to the top changed the drill order to Manufacturer → Category → Sub-Category.
- Remove — right-click → Remove Hierarchy; every field returns to the individual list.
- Save once, reuse forever — the hierarchy lives in the workbook; reopen it later and it is still there, with the data source refreshed in the background.
The instructor's workflow advice: decide with your analyst team how fields should be clubbed, build the hierarchy once, and never re-drag the same set again. "Once you get the fundamentals clear — why we group, why there is a hierarchy — these things make everything easy. It's not rocket science."
Pitfalls:
- Re-dragging the same field set field by field instead of building a hierarchy once — the hierarchy is a one-time effort reused across the workbook.
- Fear of breaking the original fields — dragging a field out of a hierarchy or removing the hierarchy leaves the individual fields intact.
- Forgetting that hierarchies persist — they live in the workbook and survive reopening, so the team's grouping decision is made once, not every session.
Recap: A hierarchy groups naturally-related fields (country → region → state → city, or year → quarter → month → day) so one drag replaces many, and drill up/down is a click. Dates come with a built-in hierarchy; custom ones are built by dragging one field onto another; they stay editable (add, reorder, remove) and persist in the workbook.
6.13 Sorting in visualizations
6.13.1 Quick sort from the axis
Once a chart exists, sorting is a one-click affair. The demo built sales by category with Sales on the label (the bars read 742k, then 719k, and so on) and sorted directly from the axis: click the axis to sort descending, click again for ascending, and a third click returns to the default order. This axis sort is the quick sort.
Worked example: sorting sales by category from the axis. The chart showed sales by category with the sales totals as labels on the bars — the largest bars read 742k, then 719k, and the rest followed. A single click on the axis sorted the bars descending, putting the biggest category first; clicking again flipped them ascending; a third click returned the bars to the default order. No menu, no dialog — the axis click cycles through the three states. Sense-check: only the bar order changes; the underlying values (742k, 719k, and the rest) stay exactly the same.
6.13.2 Sorting from the menu and from the pill
Two more routes do the same job: sorting from the top menu (the ascending/descending toolbar icons), and sorting directly from the pill on the shelf. Whichever route you take, a clear sort option always exists to reset the view to its original order.
Pitfalls:
- Forgetting that axis clicks cycle — descending, ascending, default — so a third click resets rather than re-sorts.
- Looking for a sort dialog when the quick route exists — axis click (quick sort), toolbar icons, and the pill all do it; the toolbar and pill also offer the clear-sort reset.
- Mixing this up with sorting the field list — that was the earlier distinction: field-list sorting is cosmetic navigation; this sorting reorders the marks in the visualization.
Recap: Visualizations sort in one click: from the axis (quick sort — descending, ascending, default), from the top menu's ascending/descending icons, or from the pill on the shelf. Clear sort restores the original order from any route.
6.14 Grouping data values
6.14.1 Why group
Sometimes the raw categories are too many, or you simply want to compare at a coarser level. Grouping lets you combine data values so they display together as one value. Grouping is purely for the view — the underlying data is untouched — and it gives you comparison power the raw table never had.
Intuition: Grouping is like merging drawers in a wardrobe: the clothes stay exactly where they were, but two small drawers become one bigger compartment so you can see everything at a glance. The data is untouched — only the display combines values. That is why ungrouping instantly restores the original categories.
6.14.2 Worked example: East and West
The demo used the Region dimension with Sales. With sales shown as labels, each region (East, West, Central, South) had its own bar and its own number. The requirement: "I am more interested in the combined sales of East and West." The moves: multi-select East and West, right-click, choose Group — the two regions merge into one value and the number shown is the combined sum. Ungroup restores the four separate regions.
Worked example: East + West as one bar. The Region dimension with Sales produced four bars, one per region — East, West, Central, South — each with its own sales number as a label. The requirement was to see the combined sales of East and West. Multi-selecting the East and West bars, right-clicking, and choosing Group merged them into a single value: one bar replaced two, and its label showed the combined sum of East plus West sales. The Central and South bars stayed as they were. Choosing Ungroup brought the four separate region bars back instantly. Sense-check: the combined bar's number equals East + West added together, because grouping sums the grouped values at display time rather than changing the data.
6.14.3 Creating a group from the data pane
The same operation exists in the data pane: right-click a field, Create → Group, and pick the members. The demo created the group from Region by selecting East and West, naming the group, and getting a brand-new grouped field in the data pane — usable instead of the original dimension anywhere in the workbook.
6.14.4 Editing a group
Groups stay editable. Edit Group opens the member list: add South to the East+West group and the group now covers three regions; remove members the same way; multi-select handles large member lists. If the grouping was a mistake, remove it entirely and the field returns to its original members.
6.14.5 Metros versus the rest of the country
The motivating example for grouping was metro-versus-country comparisons: select a few metro cities (Bangalore, Mumbai), club them together, and group everything else as "rest of the country" — then compare the metro total against the rest. Grouping gives you the freedom to analyze at whatever logical level you need; you are not forced to see the field values exactly as the table stored them.
Scope: Grouping is a display-time combination, not a data transformation: the underlying records stay in the source, grouped values are summed for display, and everything is reversible (Edit Group, Ungroup, or remove the group). It is the right tool when the comparison level you want is coarser than the stored categories — for example, metros versus the rest of the country — and the wrong tool if you genuinely need to change the data itself.
Recap: Grouping combines data values into one displayed value — multi-select + right-click + Group from a chart, or Create → Group from the data pane; groups stay editable (add or remove members), are view-only (data untouched), and ungrouping restores everything. The motivating pattern: compare metros like Bangalore and Mumbai against the rest of the country.
6.15 Default auto-generated fields
6.15.1 The five fields that come with every data source
Every time you bring data into Tableau, roughly five auto-generated fields appear along with your real fields. They are shown in italics — no other field is italic — which is how you recognize them:
- Measure Names
- Longitude (generated)
- Latitude (generated)
- Number of Records (the instructor also referred to it as the total record count)
- Measure Values
With a single table, the record count is per table; with multiple tables joined, each table gets its own count while the other default fields stay common to the whole data source.
Intuition: The italic fields are Tableau's housekeeping counters, like the odometer and fuel gauge of a car: you rarely read them while driving, but the moment you wonder "how many records am I actually working with?", the gauge is already there. Italics are the dashboard's way of saying "I made this field, not your source."
6.15.2 What each one shows
The demo made each field concrete with the Superstore Order table:
- Number of Records — the row count of the table: dragging it onto the label showed 9,944 — the total number of order rows.
- Measure Values — all the measure totals together in one view: the record count (9,944), discount (shown around 1,000), profit, profit ratio (which had no value in the demo), and sales at close to 2 million.
- Measure Names — the names of every numerical field in the table, so the measure values can be labeled with which measure they belong to.
A bonus behavior: double-click any measure value and Tableau automatically builds a graphic from it.
Worked example: the Order table's default fields. With the Superstore Order table as the data source, dragging Number of Records onto the label showed 9,944 — one per order row, proving the table holds 9,944 records. The Measure Values field grouped every measure total in one place: the record count (9,944), discount around 1,000, profit, profit ratio (no value in the demo), and sales at close to 2 million. Measure Names supplied the names of those numerical fields so each value could be labeled with its measure. Double-clicking any measure value built a graphic from it on the spot. Sense-check: Number of Records answers "how many rows?", Measure Values answers "what are all my totals?", and Measure Names answers "which total is which?".
6.15.3 Why bother knowing them
These fields are not something most people use daily — usually you create your own criteria and your own calculations. But they explain where the counts come from when you drag a record count or a measure value, and they show up in italics in every data pane, so recognizing them prevents confusion later.
Pitfalls:
- Mistaking the italic fields for real source columns — they are generated by Tableau, not present in your file or database.
- Expecting Measure Values to equal one number — it is a bundle of all measure totals; the record count (9,944) and sales (close to 2 million) coexist in the same field.
- Forgetting the per-table rule with joins — with multiple tables joined, each table carries its own record count while the other default fields stay common to the data source.
Recap: Every data source brings five auto-generated fields in italics — Measure Names, Longitude (generated), Latitude (generated), Number of Records, and Measure Values. Number of Records counts rows (9,944 for Order), Measure Values bundles all measure totals (sales near 2 million), Measure Names labels them, and double-clicking any measure value builds a chart automatically.
6.16 Recap and roadmap
6.16.1 What this session covered
The full checklist of what was done in this session:
- Installation — strongly recommended; the one-year student license route.
- The universal BI workflow — connect, analyze, share.
- Three connection types — file-based, server-based, saved data sources, and when to choose each.
- Live versus extract — one database icon vs two; refresh propagates or not; extracts work offline.
- Six data types — string, date, date-time, number, boolean, geographic; changeable per field.
- Data Interpreter — auto-cleaning with a review report; merged cells and PDF reading.
- Viewing data — the table view, filtering, hide/unhide fields.
- Column formatting — rename, reset name, copy a column, split, custom split.
- Sorting fields versus sorting data — the distinction the instructor explicitly flagged.
- The interface tour — menu bar, toolbar, shelves, cards, Show Me, sheets, dashboards, stories, pills.
- Dimensions versus measures — the most repeated vocabulary pair.
- Discrete versus continuous — blue pills and green pills.
- Hierarchies — the inbuilt date hierarchy and building your own.
- Sorting — from the axis, the menu, and the pill.
- Grouping — multi-select groups, data-pane groups, editing and ungrouping.
- Default fields — the five italic auto-generated fields.
6.16.2 What comes next
The next sessions go one level deeper into Tableau: joins between tables, metadata management, calculated fields (your own formulas), parameters, and sets — plus the first real, proper visualizations. The end goal is dashboarding and storytelling, and each session builds one step toward that. The closing advice repeated the central message: start using Tableau now, hands-on, so the coming sessions make sense.
Recap: The session built the complete data-side foundation of Tableau: connecting (file, server, saved sources; live vs extract), understanding (six data types, table view, Data Interpreter), preparing (rename, copy, split, sort fields), and organizing (dimensions vs measures, discrete vs continuous, hierarchies, grouping, default fields). What comes next is the analysis layer — joins, calculated fields, parameters, sets — and then real visualizations, dashboarding, and storytelling. The one action that makes all of it work: open Tableau and practice on your own data now.
Exam Guidance Summary
No exam-specific guidance was given in this session — the class is in the hands-on phase, and the instructor did not mention distributions, question patterns, or exam dates. The strongest signals about what will be expected:
- Terminology is the exam currency of this module. The instructor repeatedly said the vocabulary — shelves, cards, pills, dimensions, measures, live vs extract, hierarchy, grouping — "will help you a long way" and appears in every Tableau conversation. Expect to be asked what these terms mean and when to use each feature.
- The dimension/measure distinction is a flagged confusion point. The instructor explicitly noted that students keep interchanging these two and called the concept important, repeating the comparison several times. Know the contrast cold: descriptive vs quantitative, categorize vs quantify, never-aggregated vs auto-aggregated.
- Decision scenarios are likely questions. The lecture repeatedly posed scenario rules: file vs server (static vs real-time, sensitive vs not), live vs extract (freshness vs offline performance). These rules are natural question material.
- Interview readiness is an explicit course goal. The instructor stated that by the end of the sessions you should be able to discuss this tool confidently in an interview — so treat product knowledge and workflow knowledge as examinable.
- Hands-on skill matters. The instructor's repeated advice: install the student version and practice on personal data (expenses, portfolio), because watching videos is not enough to retain this material.
Key Industry Applications
- Tableau (Salesforce) — the product family this module is built on: Desktop, Server, Reader, Prep, Data Management, Mobile, Public, and Extensions. Tableau is a Salesforce company; classmates with Salesforce backgrounds were invited to share real-world Tableau usage.
- Power BI and Qlik — the same connect-drag-publish workflow across tools; Power BI is more mature at auto-filtering/auto-correction; Power BI appears to lack Tableau's story capability (dashboards yes, playable narratives no — verify against current versions).
- Database connectors in practice — roughly 77 built-in connectors including Databricks, MySQL, Oracle, MongoDB, Google Sheets (logs into your Google account), and SQL Server, plus JDBC/ODBC for any other database; the knowledge you need for any of them is database name, username, password.
- Sample data — the Superstore dataset (Order, People, Return tables; 9,944 order records; close to 2 million in sales) ships with Tableau as a saved data source and is the standard practice dataset.
- PDF data extraction — the Data Interpreter read an Amazon stock report PDF (historical data table) and turned it into tables, filling blank month-year values and dropping unusable lines.
- Common file formats — Excel and CSV are the everyday file connections; statistical files and PDFs are also supported.
- Learning by doing in industry — the recommended practice pattern: take personal data (expenses, investments, portfolio), build a small Excel, and drag-drop your way to understanding before touching real work.
DVI Lecture 6 notes · Tableau Desktop: Data Connection and Preparation
Sections Breakdown
The course moves from theory to hands-on work: the Tableau product family, installing Desktop with the one-year student license, how to practice, and the universal BI workflow shared by Tableau, Power BI, and Qlik.
File, server, and saved data sources; the worked Excel and MySQL connections; decision rules for choosing file versus server.
Live reads the source so refresh propagates changes; extracts are saved copies; the ZZZ-cell demo and the one-icon versus two-icon tell.
The six field types and their icons — string, date, date-time, number, boolean, geographic — and changing a field's type by right-click.
Automatic cleaning for untidy data: when it appears, the merged-header example, the review report's red and green marks, and PDF table reading.
The table view, filtering the view, and hiding and unhiding fields before building any visualization.
Renaming and resetting field names, copying columns out, and automatic versus custom split with the Order ID example.
The field list sorts A to Z for navigation only; data sorting happens at the column header — a distinction the instructor flagged explicitly.
Menu bar, toolbar, shelves, cards, Show Me, sheets, dashboards, and stories, the data pane, and pills.
Dimensions categorize and group; measures quantify and auto-aggregate — the most repeated vocabulary pair of the session.
Continuous fields show green and discrete fields blue, matching the statistics distinction; switching depends on the field type.
Grouping related fields so one drag replaces many: the built-in date hierarchy, custom hierarchies, and the location hierarchy.
Quick sort from the axis, from the menu, and from the pill; the three-click cycle and the clear-sort reset.
Combining values for display, like East and West into one bar, or metros versus the rest of the country; groups stay editable.
The five italic fields every data source brings: Measure Names, Latitude and Longitude, Number of Records, and Measure Values.
Everything the session covered, from connecting to organizing data, and what comes next: joins, calculated fields, parameters, and sets.
What the instructor's signals suggest is examinable: terminology, the dimension and measure contrast, and decision scenarios.
Tableau and Salesforce, Power BI and Qlik, the connector ecosystem, Superstore sample data, and PDF extraction.
Exam Revision Notes
Below is the distilled, exam-ready core. Every entry comes from the full explanation above. Use this section for rapid review; return to the main notes when a point needs more context.
Module 2 kickoff: from theory to hands-on Tableau
Must-know: Every BI tool (Tableau, Power BI, Qlik) shares the same three-step workflow: connect to data, drag and drop fields, save/interact/publish. Tableau Desktop is the individual analysis tool; Server shares, Reader only views, Public is the free starting point. The student license lasts one year via the student email.
⚠️ Top pitfall: Watching demo videos without installing the tool — the instructor's repeated advice is that hands-on practice on personal data (expenses, portfolio) is the only way the material sticks.
Self-check: Name the three steps of the universal BI workflow and say which Tableau product is the free starting point.
Connects to: Connecting to data: the three entry points
Connecting to data: the three entry points
Must-know: The three entry points are File, Server, and Saved data sources. Files suit static, non-sensitive, offline data; servers suit changing, sensitive, real-time data. Any server connection needs three facts: database name, username, password. JDBC/ODBC cover essentially any mainstream database.
⚠️ Top pitfall: Thinking the connection dialog knowledge transfers verbatim across databases — it does not (MySQL asks different fields than SQL Server), but the three credentials are always the same.
Self-check: A company's live order database changes every minute and holds sensitive customer data — which entry point and connection kind should you use?
Connects to: Live connections versus data extracts
Live connections versus data extracts
Must-know: Live = direct read of the source; refresh brings changes in. Extract = copy as of a time and date; refresh changes nothing. Live suits freshness and changing sources, extract suits offline work and performance. One database icon = live, two = extract; the refresh test tells you which mode you are on.
⚠️ Top pitfall: Expecting an extract to update when the source file changes — it never does; you must refresh a live connection or create a new extract.
Self-check: You see two database icons on your data source and refresh changes nothing — which mode are you in and why?
Connects to: Connecting to data: the three entry points
Data types in Tableau
Must-know: Six data types: string (ABC icon), date (calendar), date-time, number (hash icon, decimal or whole), boolean (Y/N), geographic (globe icon). Types are inferred but changeable by right-clicking the field; geographic fields carry special roles (country, state, city, postal code) that power maps.
⚠️ Top pitfall: Treating an ID column with leading zeros as a number strips the zeros; a date stored as text cannot be drilled by year/quarter/month until you change its type.
Self-check: What icon does a text field carry, and how do you change a field's data type?
Connects to: Data Interpreter: automatic data cleaning
Data Interpreter: automatic data cleaning
Must-know: The Data Interpreter cleans untidy data automatically but only appears when a discrepancy is detected; it is always optional. The review report color-codes its work: red = column headers, green = transactional data, dark green = where it made changes. PDF reading relies on it: one page, all pages, or selected page.
⚠️ Top pitfall: Expecting the interpreter to appear for clean structured data — it does not; and its guesses are not guaranteed accurate, so always review and fix the result.
Self-check: Your Excel file has merged header cells and blank year labels. What feature fixes it, and how do you know what it changed?
Connects to: Data types in Tableau
Viewing data before building
Must-know: The table view shows a real table (every field, every value) and is the inspection step before building. Filtering the view and hiding fields are cosmetic to the data: nothing is deleted, and Show Hidden Fields restores hidden fields.
⚠️ Top pitfall: Thinking hiding or filtering the view removes data — it does not; the source stays intact and everything is reversible.
Self-check: You hid a field by accident. How do you bring it back?
Connects to: Column formatting: rename, reset, copy, split
Column formatting: rename, reset, copy, split
Must-know: Rename via double-click or right-click; Reset Name restores the original. Copy a column with Control+C and paste values with Control+V. Automatic split detects the separator and creates Split 1/Split 2 columns; Custom Split takes your separator (e.g., hyphen for CA-2017-12345) and piece count. All reversible.
⚠️ Top pitfall: Trusting an automatic split to guess the right boundary for a specific format like Order ID — use Custom Split with an explicit separator and piece count.
Self-check: How do you split CA-2017-12345 into three columns, and which split type should you use?
Connects to: Sorting fields, not data
Sorting fields, not data
Must-know: Field-list sorting (data pane menu, A to Z or Z to A) is cosmetic — it reorders column names only, for navigation. Sorting data happens elsewhere: click a column header — first click ascending, second descending, third original. The instructor flagged this distinction explicitly.
⚠️ Top pitfall: Assuming sorting the field list sorted the records — it does not; only the column names are reordered.
Self-check: You clicked a column header once and the data went ascending. What do the second and third clicks do?
Connects to: Sorting in visualizations
The visual interface: a guided tour
Must-know: Shelf = where you keep fields (rows/columns); card = panels that format visuals (marks, filters, pages); pill = a field on a shelf, named for its capsule shape, blue when discrete and green when continuous. Build order: sheets → dashboards → stories. Fit options: entire view, standard, fit width, fit height.
⚠️ Top pitfall: Mixing up shelves (build the chart) and cards (style the marks); building a story before its sheets exist.
Self-check: What are the three escalating containers in Tableau, and in what order are they built?
Connects to: Discrete and continuous fields
Dimensions and measures
Must-know: Dimension: descriptive, qualitative, categorizes and groups, never aggregated, defines the bars. Measure: quantitative, numerical, auto-aggregated (sum by default; average/median/count available), defines the values. Sales by category = Category (dimension) groups, Sales (measure) supplies each bar's number.
⚠️ Top pitfall: Interchanging the two terms — the instructor explicitly flagged this as the most repeated confusion; remember descriptive vs quantitative, categorize vs quantify, never-aggregated vs auto-aggregated.
Self-check: Country names and sales figures sit in the same table. Which is the dimension and which is the measure, and why?
Connects to: Discrete and continuous fields
Discrete and continuous fields
Must-know: Continuous = unbroken scale, green pill; discrete = separate distinct values, blue pill. A share value per day is always continuous; category lists are discrete. Switching a field's nature depends on its type — a category list cannot become continuous.
⚠️ Top pitfall: Reading pill color as decoration — blue/green is Tableau's automatic signal of discrete/continuous nature and changes how the field behaves on the shelf.
Self-check: You drag a date field onto a shelf and the pill is green. What does that tell you about the field's nature?
Connects to: Hierarchies: drill up and down
Hierarchies: drill up and down
Must-know: A hierarchy = logical grouping of related fields; one drag + click-through levels. Dates auto-provide YEAR → QUARTER → MONTH → DAY. Build custom ones by dragging a field onto another and naming it; add, reorder, and remove levels anytime; Remove Hierarchy restores individual fields; hierarchies persist in the workbook.
⚠️ Top pitfall: Re-dragging the same field set repeatedly instead of building the hierarchy once — it is a one-time effort reused across the workbook.
Self-check: How do you create a custom hierarchy from Category and Sub-Category, and how do you undo it?
Connects to: Dimensions and measures, Discrete and continuous fields
Sorting in visualizations
Must-know: Quick sort = click the axis: first click descending, second ascending, third returns to default. Other routes: top menu ascending/descending icons and the pill on the shelf; clear sort resets from any route.
⚠️ Top pitfall: Forgetting the axis click cycles three states, so a third click resets rather than re-sorts.
Self-check: The bars in your chart are sorted descending. What do the second and third clicks on the axis do?
Connects to: Sorting fields, not data
Grouping data values
Must-know: Grouping combines values for display: multi-select + right-click + Group, or data pane → Create → Group. The combined value is the sum of the grouped members; Edit Group adds/removes members; Ungroup or removing the group restores the original fields. The underlying data is never changed.
⚠️ Top pitfall: Believing grouping alters the data — it is purely a display-level combination, fully reversible, with grouped values summed at view time.
Self-check: How do you combine East and West regions into one bar, and what number does the combined bar show?
Connects to: Dimensions and measures
Default auto-generated fields
Must-know: Five italic auto-generated fields: Measure Names, Longitude (generated), Latitude (generated), Number of Records, Measure Values. Number of Records = row count (9,944 for the Order table); Measure Values = all measure totals together (sales near 2 million); Measure Names labels them; double-click a measure value to auto-build a graphic. With joined tables, each table has its own record count.
⚠️ Top pitfall: Treating the italic fields as real source columns — they are generated by Tableau, and they exist in every data pane.
Self-check: Which italic field tells you how many rows your table has, and what did it show for the Order table?
Connects to: Recap and roadmap
Recap and roadmap
Must-know: The complete data-side checklist: connect (file/server/saved, live vs extract), understand (six data types, table view, Data Interpreter), prepare (rename, copy, split, sort fields), organize (dimensions vs measures, discrete vs continuous, hierarchies, grouping, default fields). Next: joins, metadata, calculated fields, parameters, sets.
⚠️ Top pitfall: Watching the sessions without opening Tableau — the closing advice repeats that hands-on practice is what makes the coming material understandable.
Self-check: List the three entry points for connecting data and the two connection modes covered in this session.
Connects to: Module 2 kickoff: from theory to hands-on Tableau
Exam Guidance Summary
Must-know: Terminology and decision rules are the examinable core: vocabulary (shelf, card, pill, dimension, measure, hierarchy, group), the dimension vs measure contrast, and the scenario rules (file vs server; live vs extract). Interview readiness is an explicit course goal.
⚠️ Top pitfall: Neglecting hands-on practice — the instructor repeatedly stressed that watching demos alone does not retain this material.
Self-check: Which vocabulary pair did the instructor explicitly flag as being interchanged by students?
Connects to: Dimensions and measures
Key Industry Applications
Must-know: Tableau is a Salesforce product; Power BI and Qlik share the connect-drag-publish workflow; connectors number about 77 plus JDBC/ODBC; Superstore (9,944 order records, near 2 million in sales) is the standard practice dataset; the Data Interpreter can extract tables from PDFs.
⚠️ Top pitfall: Over-relying on tool-specific trivia without knowing the shared BI workflow that transfers across tools.
Self-check: What three credentials does any server connection require, regardless of the database?
Connects to: Connecting to data: the three entry points, Data Interpreter: automatic data cleaning
Was this lecture useful?
BitsNotes AI Assistant
Subject Notes AssistantConfigure AI Chat
Choose how to access the chatbotSigned in as
Powered by BitsNotes — 20 messages per day. No API key needed. Want unlimited access? Use "Bring Your Own Key" mode.
Sign in to use AI Chat
Get 20 free AI messages per day to ask questions about your lecture notes. Sign in with Google or GitHub — it takes 5 seconds.
Sign In to BitsNotesSwitch to "Bring Your Own Key" tab above for unlimited access with any OpenAI-compatible provider.