Expense Tracking

You are not slow at tapping. You are slow at deciding.

Advice on categorizing expenses almost always tries to make choosing a category quicker. The bottleneck is one step earlier: the pause before you choose. Here is how to remove those pauses instead of shortening them.

The cost is the pause, not the taps

Watch someone log an expense and the slow part is not where you would expect. Typing an amount is motor memory and takes about as long as it takes. Opening the app is a second. Then the category picker appears, and something different happens: they stop.

Choosing a category is not data entry. It is a classification judgement, and judgements under ambiguity are slow in a way that has nothing to do with interface design. This has an awkward consequence for anyone trying to fix the problem by counting taps. A three-tap flow with a four-second hesitation in it is slower than an eight-tap flow with none. Most "streamlined" expense apps optimised the wrong half.

The seconds are not really the damage anyway. A pause is an exit point. Wherever you hesitate is where "I will do this later" becomes available, and later is where expense tracking goes to die - one deferred entry becomes four, four becomes a backlog, and by week three the app is a notification badge you swipe past. We wrote about that failure mode in more depth in the expense tracker guide.

So the goal is not to categorize faster. It is to have fewer categorization decisions in the first place. Those are different problems with different solutions, and only one of them is solvable.

Three tiers, in order of how much they remove

Every transaction you record falls into one of three groups: decisions you can make once and reuse forever, decisions something else can make for you, and decisions that genuinely need you. Most people treat all three as the third kind. Sorting them properly is most of the win.

Tier one

Decide once, then never again

Rent, the phone bill, insurance, the gym, four streaming services, the transit pass, the cloud storage you forgot about. For most households this is somewhere between fifteen and twenty-five transactions a month, and every one of them is the same decision, made again, with the same answer.

A rule fixes each one permanently: anything from this merchant is bills, anything from that one is transport. The work is a single audit - sit down once with a month of statements, find every line that repeats, and write the rule. It takes about twenty minutes and it retires roughly a quarter of your categorization decisions for good.

There is a side effect worth having. Nobody does this audit without finding at least one subscription they had stopped noticing, which is usually worth more than the time saved. That same list is also the raw material for your fixed-cost floor.

Tier two

Let the entry carry its own category

The fastest category is one that was already assigned by the time you looked at the screen. Three capture methods carry enough information to do that. A tap-to-pay confirmation carries a merchant name, which usually implies a category. A spoken sentence carries it in plain language - "twelve on lunch" contains both the amount and the answer. A receipt photo carries the merchant and the line items.

The structural requirement is that the parsing happens at capture, not afterwards. A voice memo you transcribe on Sunday is worse than typing on Tuesday, because it converts one small task into two and puts the second one in the future. The sentence has to be the entry.

This is the part Peggy handles without a bank connection: a phone payment becomes a categorized transaction on its own, and anything else can be spoken or photographed and comes back already labelled.

Why a wrong guess still beats no guess. People dismiss automatic categorization on accuracy, but the comparison is not "right label versus wrong label". It is "correct a wrong guess" versus "choose from a blank picker". Correcting is a one-tap decision with an anchor already on screen. Choosing is an open one. That is why a default is worth having even when it misses more often than you would like, and why the accuracy bar here is lower than it intuitively seems.

Tier three

Defer the rest on purpose

Whatever survives the first two tiers is genuinely ambiguous, and the checkout queue is the worst possible place to resolve it. You are standing up, holding things, and someone is waiting. So do not resolve it there. Send it to a single unsorted bucket, which costs zero decisions, and deal with it later.

The distinction that matters is deferring on purpose versus deferring by accident. Accidental deferral means the expense is not recorded at all, and by Thursday it never happened. Deliberate deferral means the amount and the date are safely captured and only the label is missing. A label can be reconstructed from memory a week later. An amount cannot.

Then sweep the bucket once a week, in one sitting, three minutes. Batching is genuinely faster than deciding one at a time, because you make one judgement across several similar items instead of paying the context-switch each time. And treat the bucket as a gauge: if it regularly holds more than about ten items, the problem is upstream. Tier one is missing rules, or tier two is not capturing enough.

Pre-commit the handful that are actually hard

After three tiers, the residue is small - but it absorbs nearly all the remaining time, because those are exactly the transactions that make you stop and think. The fix is to think once, now, sitting down, and never again at a till.

The mixed supermarket basket

Dominant purpose wins

Groceries plus a kettle plus batteries is groceries, because that is where most of the money went. You lose a little accuracy on the household total and you gain back a decision on the twelve baskets like it you will fill this month.

Work lunches and work travel

Categorize by who ultimately pays

Not by what was bought. If it is going on an expense claim it belongs in one "reimbursable" category regardless of whether it was a sandwich or a train, because the only question you will ask of it later is whether it came back.

Marketplaces and app stores

One default, corrected by exception

A merchant string that says "AMZN Mktp" tells you nothing about what is in the box. Default every one of them to shopping, and only go back and fix the ones large enough to distort a total. Below that, the label is noise you are paying for.

Gifts

Category by occasion, not by object

A book bought for someone else is not books. Gift spending is lumpy and seasonal, and mixing it into your ordinary categories is what makes December look like a personality change.

Anything under about €20

Never split it

The precision you gain cannot repay the decision it costs. Set the floor once, then stop reconsidering it at every till.

Underneath all five is one rule that does more work than the others combined: any consistent rule beats the correct rule applied inconsistently. The point of a category is to make this month comparable to last month. If groceries always swallows the kettle, your grocery total is slightly high and perfectly comparable. If the kettle goes to household some months and groceries in others, both totals are accurate and neither means anything.

What automation honestly cannot do

A page like this can imply that categorization becomes a solved problem. It does not, and the gaps are predictable enough to plan around.

  • Cash is invisible. No payment notification, no merchant string, nothing to parse. If a meaningful share of your spending is cash, saying it out loud is the only capture method fast enough to survive contact with a normal day.
  • First-time merchants are a coin flip. Categorization leans on merchant patterns, and a shop you have never used has no pattern. Expect the first visit to need a correction and the rest not to.
  • Payment processors hide the shop. Marketplaces, app stores and aggregators put their own name on the transaction. No system can see through that, which is why the rule above is "one default, corrected by exception".
  • Intent is unknowable. Nothing can tell that the kettle was a gift, or that this particular dinner was work. Those stay yours.

So the target is not zero decisions. It is a handful a week rather than one per transaction - which is the difference between a habit that holds and one that quietly stops in week three.

Stop choosing categories one at a time

Peggy turns phone payments into categorized expenses on their own, takes the rest by voice or receipt photo and labels it on the way in, and keeps recurring charges categorized once you have set them. What is left over is a short weekly list, not a decision at every till.

Get it on Google PlayDownload on the App Store

Frequently asked questions

What is the fastest way to categorize expenses?

The fastest category is the one you never chose. In practice that means three things in order: write a rule once for every charge that repeats, so recurring spending categorizes itself forever; use a capture method that carries the category with it, such as a tap-to-pay notification or a spoken sentence like "twelve on lunch"; and send whatever is left to a single unsorted bucket to be labelled in one batch later. Picking a category by hand, per transaction, at the till, is the slowest option available and it is the one most apps default to.

How many expense categories should I have?

Fewer than you think, if you are choosing by hand, because every extra category adds a fraction of a second of hesitation to every single entry. Seven or eight covers most people. The calculation changes if categories are assigned automatically: when the app picks, categories cost nothing at entry, so you can afford more of them. The limit then is not entry cost but attention - you only benefit from categories you actually read at the end of the month.

Why does categorizing expenses take so long?

Because it is a decision rather than data entry. Typing an amount is motor memory and takes about as long as it takes. Choosing a category is a classification judgement, and judgements under ambiguity are slow in a way that has nothing to do with how many taps the interface costs. A three-tap flow with a four-second hesitation in it is slower than an eight-tap flow with none, which is why redesigning the picker rarely helps.

Should I split a transaction across multiple categories?

Almost never, and certainly not for small amounts. A supermarket basket containing a kettle is groceries. Splitting buys you a slightly more accurate total in exchange for a decision, an interface detour and a fresh chance to abandon the entry. Set a floor - many people use something around €20 - and below it always assign the whole transaction to whatever most of the money went to. Above it, split only if the second category is one you genuinely review.

Is automatic expense categorization accurate enough to rely on?

It is more useful than its accuracy rate suggests, because the comparison is not "correct label versus wrong label". It is "correct a wrong guess" versus "choose from a blank picker". Correcting is a one-tap decision with an anchor already in place; choosing is an open decision. That makes a default worth having even when it is wrong a meaningful share of the time. What it cannot do is see cash, read a marketplace merchant string sensibly, or know that the kettle was a gift.

How do I handle expenses I cannot categorize immediately?

Put them somewhere on purpose instead of leaving them uncategorized by accident. An explicit unsorted bucket captures the amount and date in zero decisions, and a label can be reconstructed from memory a week later. An amount cannot. Sweep the bucket once a week in one sitting, which is faster than deciding one at a time because you are batching similar judgements rather than context switching. If the bucket regularly holds more than about ten items, the problem is upstream: your rules and your capture method need work.

Explore Peggy guides and comparisons