About Puzzleshooting

Puzzleshooting is a series of realistic data troubleshooting scenarios — the kind of messy, half-documented problems that actually show up in production, not textbook exercises. Each puzzle drops you into a situation: bad data, a broken pipeline, a script that worked yesterday and doesn't today. You get the context, the files, and the constraints. You don't get the answer until you've had a shot at it.

This isn't about knowing the "right" tool. Most puzzles can be solved in various ways: with a shell one-liner, a Python script, SQL, or whatever you're fastest in — the point is the diagnostic process, not the syntax. That said, sometimes there are constraints designed to help you learn new techniques.

Who this is for

Anyone who works with data and wants to get sharper at the actual job: reading messy inputs, forming a hypothesis, checking it, and fixing it cleanly. Useful whether you're a few years into a data engineering role or self-taught and want reps on real-shaped problems.

How to use it

Each puzzle has a problem page with the scenario and sample data, and a separate solution page with a full writeup — walk through it, try your own approach, then compare notes. New puzzles land regularly; subscribe below to get them as they publish, along with the occasional deeper writeup that doesn't fit the puzzle format.

Start Here

New here? Start with the archive and pick whatever looks interesting — there's no required order though certain puzzles may build on others.

Short version: read the scenario, try to solve it with the suggested tools, then check the solution page to see how it compares. That's it.