I was the tax office’s unpaid clerk, until I delegated it to an AI

Every summer, a shoebox wins.
For about fifteen years, preparing my German tax return has cost me a weekend plus a few evenings, and quite a lot of swearing in between. Not the tax return itself: I have a Steuerberaterin for that, and she is excellent. What costs me the weekend is everything that has to happen before she can start. Remembering which receipts exist. Finding them. Naming them. Sorting them into the right folders. Building the spreadsheets that explain the numbers, and the notes that explain the spreadsheets.
This year I handed most of that to an AI. It still took a couple of days. But for the first time since I have been filing, it didn’t suck, and the result was better than anything I had produced on my own.
The work has a name, and it isn’t complexity
I always assumed the problem was that German tax law is uniquely, insanely complicated. That turns out to be one of those things everybody knows and nobody has checked.
For two decades it circulated in German public debate that 60 to 80 percent of the world’s tax literature is written in German, the standard shorthand for “we have the most complicated tax system on earth.” However, in 2004 a tax professor named Albert J. Rädler went to the International Bureau of Fiscal Documentation in Amsterdam, the largest tax library in the world, with a tape measure, and measured the shelves. The German share was around ten percent. Later database counts put it between 15 and 22 percent. English-language material sat at 54 to 65 percent.
The measured picture agrees with the tape measure. In the Tax Complexity Index 2024, Germany scores 0.40 on a scale that runs from 0 to 1, against a global average of 0.38 across the 71 countries surveyed. The United States sits at 0.39, the Netherlands at 0.33. Brazil tops the table at 0.56, Estonia anchors the other end at 0.17. So Germany is a little above the middle, and that is the whole story. The index authors put it plainly themselves: countries “whose tax systems are often considered the most complex have a medium overall level of complexity.” (Note: that index measures corporate tax complexity as perceived by advisers to multinationals. I haven’t found a comparable ranking for private individuals, yet.)
So if complexity isn’t the enemy we thought it to be, why does it feel like one?
Because complexity is not the real issue here, it’s in who does the assembly.
Positions in this image are illustrative rather than plotted. The horizontal axis comes from the complexity index; the vertical one is my own measure, built from what each tax authority publishes about its own process. In Norway, if you do nothing, the pre-filled return is considered submitted. In Finland: “If all the information is correct, you do not need to do anything.” In Denmark, 89 percent of the numbers on your assessment arrive from third parties before you see it. Sweden lets you confirm by text message. Estonia pre-fills income and deductions, and calls it a three-minute job. In the UK and Japan, most employees never file at all.
Germany sits in one corner with the United States. The kicker: the German state already has my data. Its digital tax portal ELSTER offers a Belegabruf, which hands me back my wage statements, my pension notifications, my health insurance contributions. What it does not help me with is deductions. No Werbungskosten, no haushaltsnahe Dienstleistungen, no donations, no rental income. Exactly the side that produces the refund, and exactly the side that consumes the effort.
The state only does the “income” part of its job.
Ivan Illich coined the word shadow work for this in 1981. The unpaid, invisible labour a system requires from you before it will deliver anything. Commuting. Self-checkout. Filling in forms. Cal Newport and others later extended it to knowledge work: the sorting, tagging and deciding that has to happen before any visible work can begin. It is typically a large share of the effort, and it is never the part anybody plans for.
That’s what my August weekend really was. I wasn’t doing my taxes. I was working an unpaid shift as a clerk for the Bavarian tax administration.
This sounds exactly like the kind of work that should be automated away. So I asked myself: can AI absorb my shadow tax work?
What I delegated, and what I couldn’t
I used Claude Cowork with the Opus 5 model, and my rule was simple: delegate everything except the parts that need hands. Claude has no legs, so fetching paper receipts from drawers and running them through the scanner stayed with me.
I started with three things: last year’s completed tax folder, the checklist I have refined over a decade, and this year’s folder, which was a mess. Then I wrote a delegation brief, which came down to roughly this:
Hey Claude, here are two folders with my tax documents, one for last year, and one this year. Here’s also my checklist that I used last year. Can you help me prepare my tax receipts and build the spreadsheets for this year’s taxes? Let’s use the folder structure and my checklist as a starting point for what we need to do here. Let me know when you have any questions or anything is missing and let’s work through this together.
The last part did more work than the rest combined. It turned the job from “produce an answer” into “run a process and tell me where it stalls,” which is a much better fit for a task where I am the one holding all the missing pieces.
Over the following week, Claude sorted receipts into folders, built spreadsheets for the rental property, the freelance side, the book royalties (I am a co-author of a German Paleo cookbook), and the share sales, wrote explanatory notes for each, and produced a two-page cover document so my adviser could see how the folder was built. Oh, and it also did the same for my father’s taxes (though those were much simpler) as well. While it worked, I did other things. That is the part that actually changed the experience: not that it was faster, but that it became a series of short, in-between tasks that allowed me to do other work at the same time.
The key thing that made it work: make it checkable
The obvious objection to all of this is: large language models are unreliable at arithmetic, they state wrong things with total confidence, and tax numbers have consequences. Why on earth would you trust one with this?
You shouldn’t. But humans are bad at arithmetic and they’re fallible, too. So the answer is tools and structure. Here’s what worked surprisingly well for me:
Never let it do mental arithmetic. Every number in every table is a formula in a spreadsheet, or the output of a script I can read and rerun. The model’s job is to decide what to compute, not to compute it. This alone removes most of the risk people worry about.
Build in control rows that must compute to zero. Each spreadsheet ends with a reconciliation: the sum of the individual receipts minus the total from the bank statement, the declared proceeds minus the actual money that arrived. If a control row is not exactly 0,00 €, something is wrong and it is visible immediately, without anyone having to notice. This caught a real error, too: I had a sign flipped in a reconciliation of four incoming transfers, and the control row refused to close until it was fixed. It’s like unit testing, for spreadsheets.
Verify against the source, mechanically. When a table is derived from a document, don’t proofread it. Write a small script that pulls the numbers out of the original PDF and compares them, line by line, to what ended up in the table. It takes two minutes and it checks fourteen positions more reliably than I ever will at eleven at night.
None of this makes the output correct. It makes wrongness visible, which is what turns it into a run-test-fix loop. AI agents are very good at those.
Small meta-moment: while researching the tax complexity numbers, Claude and I kept arriving at different rankings for Germany. Claude had sorted the countries from most complex downward and got 31 of 71. The website counts from the least complex upward and shows 41. Same position, counted from opposite ends of the same list, and indeed 71 minus 41 plus 1 is 31. Neither of us invented a number. One of us silently reversed a sort order, and it only became visible when we put the two sources side by side. Which is exactly the argument for control rows.
What worked well
It was more accurate than I had been. Because it had last year’s filed return and the resulting assessment from the tax office, Claude could compare my old spreadsheets against what my adviser had actually submitted. It found places where I had been wrong for years. For example: I had been halving my adviser’s fee using the net amount, where she correctly uses the gross. Small money, and she fixes it anyway, but I had made the same mistake every year without noticing.
It answered its own questions. At one point the open-questions list for my adviser had grown to around thirty items. Far too many for a one-hour appointment, and frankly a bit rude. So instead of answering them myself, I handed over last year’s complete filed return and the tax office’s assessment, and asked Claude to answer as many as it could from precedent. It got the list down to seven, all of which were genuinely new situations from 2025. The lesson: feed the previous outcome back in and let the questions answer themselves. It’s the single most valuable thing I learned all week.
It reviewed its own work. Something I do with Claude Code on software projects all the time: code reviews. “Why not for tax work, too?” I thought. So I asked Claude to review every receipt, spreadsheet and note again, as a final step. It found a few things here and there that it was able to fix and simplify, increasing my confidence in the results. If you want to go one step further: open a fresh session and ask for a review with empty context, giving you fresh eyes on top of the mere self-review I let Claude do.
It went shopping. The last test was giving it access to my browser and asking it to work through my 2025 Amazon order history. It scanned three hundred orders, proposed a shortlist for me to approve, then downloaded the invoices we agreed on and filed them, correctly separated into employee expenses and freelance expenses. I hadn’t even thought about splitting those two yet.
A short detour for nerds
Skip this section if shell scripts aren’t your idea of a good time, but this made me smile:
Some of my receipts are scans: image-only PDFs with no text layer. I noticed Claude struggling with a couple of them and asked what was going on. It explained to me that it was trying to avoid looking at too many images to preserve context size and after I offered to help convert those into smaller or more efficient formats, instead, it went looking, found that I already had the Tesseract OCR engine installed, checked that the German language data was present, and wrote itself a small conversion script to make those PDFs scannable and searchable.
What I find interesting there isn’t the OCR. It’s the shape of the move: the agent hit a capability gap, inspected the environment it was standing in, and built the missing tool for itself instead of escalating to me. That is a different thing from “answering well,” and it is the behaviour that made the week feel like working with a colleague rather than operating a machine.
Where it went wrong
Of course there were also some bumps along the way, so here they are:
It was too diligent. The thirty-question list wasn’t a one-off. Left alone, it would happily add every ambiguity it encountered to the pile. Reminding it to curate that list down was my job, as was the act of guiding and overseeing the whole process.
A silent failure nearly slipped through. During the Amazon run, sixteen invoice downloads reported success and none of them arrived on disk. Chrome blocks multiple automatic downloads per site by default, and the block is invisible to the thing doing the downloading. Everything said “200 OK”. Nothing existed. However, Claude figured this out by itself and asked me to switch multi-downloads on in the Chrome preferences for the amazon.de website. A setting I didn’t even know existed.
And my favourite: To separate private Amazon orders from deductible ones, Claude wrote a keyword filter. One of the private-purchase keywords was hund, German for dog, because we own a dachshund who accounts for a fair share of our Amazon orders. That pattern promptly matched the word Jahrhunderts (century), and quietly excluded a book I actually wanted to claim. It caught its own bug on review and reported it to me. I have made worse mistakes with regular expressions.
Not a lot of catastrophes, then, and most of these, Claude actually fixed itself, with some nudging.
What I would not hand over
I gave an AI my salary, my share transactions, my bank statements and, because I also prepare my father’s return, someone else’s financial data as well. That deserves some consideration.
Two rules I observed here: Everything ran inside my own folders and a sandbox tied to my account, and the finished package went to my adviser in a password-protected archive rather than by email. Nothing was published, sent or submitted anywhere without me clicking the button. And for my father’s documents I applied a simple rule: data that isn’t mine gets the same handling I would want for my own, which in practice meant the same folder, the same encryption, and no shortcuts.
I know this sounds more relaxed than others might be comfortable with. However, it’s the same private life many of us hand over to social media, search engines, and “free” services (which never really are). In my case, I prefer vendors I vet and pay for, so I have at least some legal basis for working with them.
I also kept the receipt gathering mostly manual, and that was deliberate rather than timid. A braver version of me would have pointed the browser integration at every merchant account I own. But collecting is the step that introduces a natural friction that makes me stop, review, and consider my next move, keeping me in tight lock-step with the overall process I’m ultimately responsible for.
The bottleneck moved
The whole thing still took about as long in wall-clock terms as it always does.
For a while that bothered me, and then it stopped, because it isn’t a failure. It’s Goldratt’s lesson from The Goal. When you speed up one station in a chain, the constraint doesn’t disappear, it moves to the next slowest station. I removed the sorting, the spreadsheet building and the note writing. What was left standing was walking to the drawer, feeding the scanner, and making the judgement calls only I can make. That is now my bottleneck, and it is a far better bottleneck than the one I had.
The same logic scales up, which is why I don’t think this makes tax advisers obsolete. Quite the opposite. An AI cannot be held liable for its work, but my adviser can. An AI is not a regulated profession, hers is. Collecting and pre-processing documents was never the valuable part of what she does, it was the tedious part, and it was mostly being pushed downhill onto me anyway. If that layer gets absorbed, what is left is judgement, representation and responsibility, which is the part I am actually paying for.
This pattern goes well beyond tax. Insurance claims. Permit applications. School forms. Paperwork. Anywhere an institution holds the data and still makes you do the assembly, there is shadow work waiting to be handed off. Estonia closed that gap fifteen years ago by fixing the system. Most of us can’t wait for our own governments to do the same. The good news is that we no longer have to.
Some suggestions for you to try out
If you want to test this on your own pile, start smaller than I did and use these three habits: formulas instead of mental LLM arithmetic, validation and testing mechanisms, and a mechanical check back against the source. Then pick one bounded, genuinely tedious assembly job, hand it over, and see what breaks. You will learn more from the first thing that goes wrong than from the first thing that goes right.
My return is with my adviser now. I am curious what she makes of the folder, and slightly nervous, the way you are when you have shown your work for the first time.
What’s your shadow work? The unpaid shift you do repeatedly for a system that already has your data? Let me know!
