Ask ten business owners whether AI is saving them time and ten will say yes. Ask them how much and you get a number that arrived by vibes.

"Probably ten hours a week." Based on what? Based on the feeling of having done something faster. Which is a real feeling and also completely unreliable, because the thing you actually notice is the task that got faster, and the thing you never notice is the two new tasks that appeared to feed the tool.

I got tired of not knowing, so I spent four weeks measuring it properly. This is the tool I used, how I set it up, what the test looks like, and the number I got, which was not the number I expected.

The tool

Rize is a time tracker that runs in the background on your machine and categorizes what you are actually doing rather than what you told it you would be doing. No timers to start. No projects to select. It watches which application has focus, for how long, and buckets it.

That last part is the whole reason it works. Every time tracker that requires you to press start fails within nine days, because the moment you are busy enough for the data to matter is the moment you stop remembering to press start. Passive collection is the only kind that survives contact with a real week.

You can find it at rize.io.

What it actually does well

Three things, and I want to be specific rather than list features.

Category accuracy out of the box is genuinely good. It knows a terminal is work and a video site is not without you configuring anything. More usefully, it distinguishes between categories of work. Communication versus deep work versus meetings versus admin. That distinction is the thing that makes the AI question answerable, because AI does not reduce your hours, it moves them between categories.

Focus session detection is the sleeper feature. It identifies uninterrupted stretches and tells you how long your longest one was and how many you had. I care about this more than total hours now. A day with four hours of work in six fragments is a worse day than three hours in two blocks, and until I had the number I did not know how badly fragmented I actually was.

The daily summary is short enough to read. This sounds trivial. It is not. Most analytics tools produce a dashboard that requires a decision to open, and a dashboard requiring a decision does not get opened. A short summary that shows up and can be read in thirty seconds gets read.

What it does not do well

It cannot see inside the browser well enough for knowledge work. Everything is a browser tab now. Research, writing, the CRM, the ad platform, the model you are prompting. Rize sees "browser" and does a decent job with domains, but if you spend six hours a day in fifteen tabs across four jobs, the categorization gets mushy and you will spend the first week correcting it.

No native mobile picture. If a meaningful part of your work happens on a phone, that time is invisible, and the report will quietly overstate how much of your day was deep work because the fragmented phone hours are not in the denominator.

The manual correction loop takes real effort up front. Budget an hour across the first week reclassifying things. It learns, and by week two it is mostly right, but anybody who tells you passive tracking is zero setup is selling something.

The four week test

Here is the actual protocol. This is the part worth keeping even if you use a different tool.

Week one: baseline, change nothing.

Install it. Correct the categories as they come up. Do not change how you work. Do not optimize anything. You are collecting the before picture and the temptation to start improving immediately will ruin the whole experiment.

At the end of week one, write down four numbers. Total tracked hours. Hours in deep work. Hours in communication and admin. Number of focus sessions over forty five minutes.

Week two: baseline again, still change nothing.

One week is not a baseline, it is an anecdote. Week two exists to tell you how much natural variance there is between weeks, and there is more than you think. If deep work is eleven hours in week one and nineteen in week two, then a two hour improvement in week four means nothing, because two hours is inside the noise.

Most people skip this week. Most people therefore cannot tell improvement from Tuesday.

Week three: introduce one change.

One. Not a stack overhaul. One workflow moves to AI, or one existing AI workflow gets rebuilt.

Write down what you expect to happen before you start. Actually write it. "I expect this saves me about four hours a week on proposal writing." You want the prediction on record because the gap between what you predicted and what happened is the most educational number in the whole exercise.

Week four: same change, measure.

Let it run one more week so you are not measuring the learning curve. Then compare week four against the average of weeks one and two.

What I found, and it was annoying

My prediction going in was that AI content workflows were saving me somewhere around eight hours a week.

The measured number was 5.5 hours of reduction in writing and drafting time. So far so good, roughly in the range, though lower than I thought.

Then I looked at the other categories. Time in what Rize buckets as tool configuration and admin went up by 2.75 hours a week compared to baseline. Prompt fiddling. Fixing a scenario that broke. Reviewing output. Reorganizing folders. Reading about a new model to decide whether to switch.

Net was about 2.75 hours. Not eight. Under three.

That is still real. Three hours a week is a hundred and fifty hours a year and I will take it. But it is roughly a third of what I would have told you at a dinner, and I would have said the bigger number with total confidence, because the five and a half hours I saved were visible and vivid and the two and three quarter hours I spent were spread across forty small invisible moments.

FROM THE AI NEWSROOM

The AI Workflow Blueprint

The exact systems behind everything in this issue. Routing tables, gate logic, the scenario templates, and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.

Get the Blueprint for $47

.....

The second number, which matters more

The hours were the boring finding. Here is the interesting one.

My focus sessions over forty five minutes went from four a week at baseline to nine a week in week four. More than doubled.

That is the real effect and it has almost nothing to do with time saved. What changed is that AI removed the specific kind of task that used to break my day into pieces. The forty minute "just knock out this draft" jobs that sat between meetings and made a two hour gap useless for anything real. Those got compressed into ten minutes, which meant the gaps became usable for actual work.

Same total hours, restructured into bigger blocks. That is worth more than the three hours, and I would never have known it happened without the data, because there is no feeling associated with "my day had fewer fragments."

Which suggests most people are measuring the wrong thing. Do not ask whether AI saved you hours. Ask whether it gave you back contiguous time. That is the number that determines whether you can do work that requires holding something complicated in your head.

The setup that makes the data usable

A few configuration notes that saved me from bad data.

Split your AI time into two categories, not one. "AI production" for time where a model is doing work for you, and "AI overhead" for time spent configuring, debugging, prompting, and reading about tools. If you lump them, you will never see the finding above, and the finding above is the whole point.

Tag meetings separately from everything else. Meeting hours dominate most weeks and they swamp the signal. You want to be able to look at your non meeting hours in isolation, otherwise a week with three extra calls looks identical to a week where your workflows broke.

Do not track the tracking. Some people get very into this and start spending twenty minutes a day looking at their own dashboards. That is a new job you have invented for yourself. Look at it once a week, on the same day, for ten minutes.

Export weekly to a sheet. The in app history is fine but you want your own copy for the comparison, and you want it in a format you can put a formula on. I pull the weekly export into a Google Sheet and average the baseline weeks in a cell, because doing that arithmetic by hand is exactly the kind of small friction that ends experiments.

Who should not bother

If your week is over sixty percent meetings, this will tell you what you already know and there is not much to optimize. Go fix the calendar first.

If you work primarily on a phone or across three machines, the picture will be incomplete enough to mislead you, and an incomplete picture is worse than no picture because you will trust it.

And if you are not going to run the full four weeks, do not start. A one week snapshot with no baseline is a number that feels like evidence and is not, and you will make a decision on it.

The broader point

There is a lot of noise right now about how much AI is or is not delivering. Some of it is genuinely useful analysis and some of it is people with a position to defend. What almost none of it is, is a measurement of your business.

The honest answer for most operators is that they have no idea what their AI stack is doing to their week, in either direction. They have a feeling, and the feeling is shaped by the last vivid thing that happened.

Four weeks and a passive tracker gets you an actual number. The number will probably be smaller than your guess and the second order effect will probably be bigger than you expected. Both of those are worth knowing, because the first one stops you overinvesting and the second one tells you what to optimize for.

If you run this and the number comes back near zero, that is not a failure of the experiment. That is four weeks well spent, because you just learned that the thing you have been telling yourself is not true, and you can go find where the time actually is.

FROM THE AI NEWSROOM

The AI Business Accelerator

For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.

Join the Accelerator for $97

.....

The one week version

If four weeks is more discipline than you have right now, do the cheap version.

For the next five working days, at the end of each day, write down two numbers in a note. How many hours felt like real work. How many uninterrupted stretches over forty five minutes you got.

That is it. Ninety seconds a day, from memory.

It is less accurate than passive tracking and it is dramatically better than nothing, because by Friday you will have five data points and a pattern, and the pattern is usually enough to tell you which day of your week is broken.

Jordan

The AI Newsroom | Practical AI for people with a business to run.