Data engineering interviews lean heavily on SQL and Python. Expect to write queries by hand, explain how indexes and transactions work, and discuss how data is modelled and partitioned at scale.
INNER JOIN returns only rows that have a match in both tables, while LEFT JOIN returns every row from the left table and fills the right table's columns with NULL when there is no match.
WHERE filters individual rows before grouping, while HAVING filters groups after GROUP BY, so only HAVING can use aggregate functions like COUNT or SUM.
DELETE removes selected rows and can be rolled back, TRUNCATE quickly removes all rows and keeps the table structure, and DROP removes the entire table including its structure.
A primary key uniquely identifies each row in its own table and cannot be NULL, while a foreign key is a column that references the primary key of another table to link the two.
A list is mutable and can be changed after creation, while a tuple is immutable; tuples are slightly faster, use less memory and can be used as dictionary keys when their items are hashable.
Git is a distributed version control system that tracks changes to files on your machine; GitHub is a hosting service for Git repositories that adds collaboration features such as pull requests, code review and CI.
git fetch downloads new commits from the remote without changing your working branch; git pull runs a fetch and then merges (or rebases) those commits into your current branch.
Give a 60-90 second summary in three parts: what you do now, one or two achievements that show your strengths, and why this role is the logical next step.
Connect two or three of your strongest, provable skills directly to the needs in the job description, give a short piece of evidence for each, and show genuine interest in the team鈥檚 work.
*args collects extra positional arguments into a tuple and **kwargs collects extra keyword arguments into a dictionary, which lets a function accept a variable number of arguments.
Use a subquery that takes the MAX salary below the overall MAX, or use DENSE_RANK() and pick rank 2, which also handles ties and generalises to the Nth highest.
Find duplicates with GROUP BY and HAVING COUNT(*) > 1, then delete them by numbering rows within each duplicate group with ROW_NUMBER() and removing every row after the first.
Normalization is organising tables to reduce redundancy and update anomalies: 1NF means atomic values, 2NF removes partial dependencies on part of a composite key, and 3NF removes dependencies between non-key columns.
Default argument values are evaluated once when the function is defined, so a default list or dictionary is shared between every call that does not pass its own, causing data to leak between calls.
A shallow copy creates a new container but reuses the same nested objects, while a deep copy recursively copies everything so the new object is fully independent.
A classmethod receives the class as its first argument (cls) and can create or configure instances, while a staticmethod receives no automatic argument and is just a function placed inside the class namespace.