Lesson 1  ·  Oct 4, 2026
Data Engineering Fundamentals

Reading data with SELECT and WHERE

The first thing you do with any new table is look at it: pick the columns you care about, keep the rows you care about. In SQL that's two clauses, and they map exactly onto pandas.

The code

Say we have a customers table with name, age, and city columns:

SELECT name, age
FROM customers
WHERE age > 30;

Each clause has one job:

In practice

Every data task starts here. Before any model, dashboard, or pipeline, someone writes a query that reads a table and filters it down to the relevant slice. Getting fast at reading SELECT ... FROM ... WHERE ... is getting fast at the first 30 seconds of every analysis.

Python bridge: WHERE is pandas' boolean filter, SELECT is the column list. The query above is customers[customers.age > 30][['name', 'age']]. If you can write the pandas, you already understand the SQL — it's just the vocabulary that's new.

Try it

Practice question: before running anything, predict what this returns — then change the filter to age < 25 and predict again:

SELECT name, city
FROM customers
WHERE age > 30;
Show answer
SELECT name, city
FROM customers
WHERE age > 30;

Returns the name and city columns — but only for customers older than 30. Same two columns as the SELECT lists, fewer rows than the full table.

SELECT name, city
FROM customers
WHERE age < 25;

Still name and city, now for customers younger than 25. The SELECT never changed, so the columns stay the same — only the WHERE filter changed, so a different set of rows comes back.

← All posts