Aug 20, 2026 · first shared on LinkedIn
I stopped grading my AI agents on tasks completed
I stopped counting tasks. I started counting projects. For months my AI agents got graded on tasks completed, and the days looked great on paper. Lots of little things done, checked off, gone. But this week the unit of measurement for everything I run changed, and I think the reason applies to any business, with or without AI in it.
Busy is not the same as delivered
A task list rewards motion. Every checked box feels like progress, so the setup naturally fills the day with things that are easy to check off. But like you know, a pile of finished tasks can still add up to nothing a customer ever sees. You can complete a hundred tasks around a product and never actually ship the product. Nobody plans it that way. It is just what happens when the scoreboard counts activity instead of outcomes.
So the scoreboard changed. My agents are now graded on projects delivered, not tasks completed. Same work underneath, completely different behavior on top, because whatever you measure is what the whole system quietly starts optimizing for.
What counts as delivered?
Ready for the market, not polished. That is the definition, and it is deliberately uncomfortable. If someone outside could use the thing today, it is delivered. If it still needs one more round of tweaks that only I would notice, it is not more delivered, it is just more polished. Grading on delivery pushes everything toward the moment a real person can touch it, instead of toward another comfortable week of internal improvements.
Why rank the projects?
Three projects are live right now, and they sit in a ranked list where the one on top is the one that matters most. The position is the priority. That sounds almost too simple, but it removes a real cost: the daily re-negotiation. When priorities live in your head, every morning starts with a small silent debate about what matters today, and the loudest or newest thing usually wins. When the priority is a position in a list, the debate is already settled. You change the order deliberately or you follow it.
It also forces goal-level visibility. A task list tells you what everyone is doing. A ranked project list tells you what the business is actually trying to deliver next. Those are very different questions, and only the second one tells you if the weeks are adding up to anything.
What happens when a project is too big?
This is the rule I like most. Anything that would take more than 7 days gets flagged as too big and broken into milestones. Because a project with no end in sight is where work goes to hide. It is always “in progress”, it never fails, and it never finishes. Splitting it into milestones means something must be delivered inside the week, so a stalled project gets caught in days, not months.
Think of it as counting kilos lost instead of gym sessions attended. Sessions measure effort. Kilos measure whether the effort changed anything.
Can you do this without AI?
Completely. Nothing about this rule needs software. Take what you are working on, group it into actual projects, rank them so the top one is unmistakable, define delivered as ready for a customer, and split anything that would run past a week. The AI agents in my setup just enforce it without getting tired, the same way they enforce every other rule I used to break myself.
Tasks measure motion. Projects measure whether anything shipped. Change what you count and watch what changes.
I build AI agents that run all parts of my online business. Check them out at pelleforsman.com.