
Why Some AI Workflows Stick—And Others Don’t
These four questions helped me decide which ones to keep, redesign, revisit, or retire
Was this newsletter forwarded to you? Sign up to get it in your inbox.
I was on a video call with Every CEO Dan Shipper, getting a tour of the increasingly elaborate ways he uses OpenAI’s Codex app to run his work life, when I saw the next AI workflow that was going to change mine. At least, that’s what I told myself.
Dan was sharing an early version of Tend, the Codex-native system he created to pull email, Slack, meeting notes, and company updates into one place and help you act on them.
The moment the call ended, I gave the transcript to Codex and had it build my own version of the tool, which, lacking Dan’s flair for naming things, it called Attention Desk. I could picture the person who would use it: me, but better. I would spend most of the day in deep work, surface a few times to see who and what needed me, and disappear again—responsive, efficient, and 10 times more productive than I’d been before.
As of this writing, Attention Desk sits abandoned in my pinned chats, the little blue dot beside its name the only reminder of how sure I was that it would transform my work.
This seems to be a pattern in the way I use AI: I see an impressive demo or have one surprisingly good session with a model and decide to reorganize my work around it. Then the novelty wears off. The next thing I know, weeks have gone by and I haven’t touched the workflow I was sure would change everything. There’s a veritable Island of Misfit Workflows adrift on my desktop, from app ideas that collapsed under the weight of their own maintenance to AI skills I built and promptly abandoned. And yet writing with my compound writing plugin or navigating my workday with help from Codex already feels like I’ve been working this way for years.
I wanted to understand why some AI workflows become essential while others have the staying power of a celebrity romance. So I did a post-mortem on the chief offenders and looked at relevant behavioral science research. It turns out the difference comes down to what a workflow asks of my time, energy, and sanity—and what it offers in return.
An abandoned workflow is not a character test
Once I noticed that little blue dot next to Attention Desk and the fact that I wasn’t responding to it, I let it sit like a check-engine light I’d decided to live with. Over time, the notification started to feel like an indictment of my character.
My first reaction was to label it a me problem. I simply lacked the stick-to-itiveness or sophistication to get these systems to work. Somewhere out there, other, better people were adopting these workflows and living their 10-times-more-productive lives while the permanent underclass had a space with my name on it.
Maybe that’s just my own unique brand of catastrophizing, but I suspect this feeling of wanting to build all the workflows—particularly the workflows of people we perceive as successful—is common. Model capabilities and the best practices for working with them are changing so fast that it’s understandable to want to follow the lead of people who are winning at AI. When you try what works for them and don’t get their results, it’s tempting to assume you’re doomed to failure. At least, that’s what happens if you’re me.
Rationally, I know that “either I can make this workflow work or I’m a failure” is classic black-and-white thinking. A workflow may go unused for a boatload of reasons: The problem is too infrequent, the trigger doesn’t go off, or the output creates more work than it saves. A tool can fail because my habits don’t support it; my habits can also be rational responses to the work I actually do.
I built Attention Desk for a workday I don’t have
I wanted what Tend offered: to sink into deep work without wondering what was happening in Slack, then surface to find nothing had caught fire. But Dan is a busy CEO pulled in enough directions that Tend’s ability to gather his communication gives him time and mental space back. I get an average of six to eight inbound Slack messages a day, which I can manage myself without much stress. For me, Tend proved to be a solution to a problem I simply didn’t have.
With my Tend dupe, I was also building a workflow that would compete with my own deeply ingrained behavior. I have an anxious attachment style when it comes to work, and checking Slack is my favorite sanctioned procrastination activity. I reliably check Slack 40 to 50 times per day—and every time, I feel relief from the discomfort of writing, or the motivating hit of a fresh news drop. Attention Desk’s attempt to deliver peace of mind couldn’t compete with those rewards.
Behavioral psychology lore says I should have given Attention Desk 21 days to stick as a habit, but that data point turns out to be a myth. A 2024 literature review found that the few relevant studies showed habit adoption medians around two months and ranges from four to 335 days. Three days may not have been long enough to create a habit, but it was plenty of time to discover I didn’t want to spend weeks or months forming this one.
The research also has ideas about how I could force an AI workflow to stick. For instance, I could make an explicit plan: If I am about to check Slack, then I will open Attention Desk. A 2024 meta-analysis covering 642 tests found that clear plans grounded in the specific context where the behavior happens can help, especially when motivation is already strong. There’s the catch: An if–then plan can help me act on a goal that feels important and compelling. It cannot make me need a CEO dashboard or want to stop using Slack as an emotional-support procrastination device.
Attention Desk wasn’t defeated by my failure to perform the right behavior-design incantation. In hindsight I can see that I made a rational choice to abandon the tool; it was built for someone else’s problem and pitted against a behavior that solved my own.
I kept building tools for a person I wasn’t
Attention Desk has neighbors on the island, each built for a version of me that doesn’t exist. I set up a Codex automation to generate daily X and LinkedIn post recommendations based on my work. I pictured myself effortlessly hitting the three-times-a-week activity cadence I saw recommended for LinkedIn.
Reader, I have never shared one of those posts.
In the cold light of day, I realized I’m a spontaneous, shoot-from-the-hip kind of social media user. When I have something to say, I’ll say it. Suggested content gives me another queue to fall behind on, and I emphatically do not want that. For me, the annoyance outweighs the value of posting more. Becoming an X poster extraordinaire will have to wait.
In the 1980s, psychologists Robert Wicklund and Peter Gollwitzer described a phenomenon uncomfortably close to my pattern of optimistically building tools for an aspirational version of myself: “symbolic self-completion.” People who fall short of an identity they care about can reach for its symbols, including tools and skills. Their experiments did not test AI workflows, but still—building the machine that would make me a disciplined public thinker felt remarkably similar, for an afternoon, to being one.
I doubt I’m alone in carrying around these more motivated selves. They’re aspirational, like a workout outfit you buy optimistically. One takeaway from my exploration of these failed projects could be to embrace the Katie I am, not the one I hope to be tomorrow.
A workflow can be worth building and not worth keeping
Longtime Working Overtime readers will remember Margot, the AI assistant I built to manage my other workflows. I set her up after seeing what people like Nat Friedman and Claire Vo were doing with OpenClaw on an Every camp on the topic. I imagined an AI companion fine-tuned to my needs could supercharge my day-to-day life.
Life with Margot was great for the first several weeks. After the “it’s alive” moment and eager exploration of her Google Calendar navigation, she became more trouble than she was worth. For example, every time I tried to change the model powering Margot, she quit working altogether and I had to revive her. A month in, I realized I was spending more time maintaining Margot than I got back.
I left Margot behind, but her legacy carries on. Architecting her memory taught me how important it is for agents to find the right context. That lesson fed into my current desktop setup, where tailored context helps Codex or Claude find what they need.
Living as we do in the Wild West days of AI, we’re all being tasked with figuring out what is and is not worth the effort to set up and maintain. There’s value to tinkering with these temperamental, wonky workflows because it makes us savvier users of the models and apps we work with everyday.
The workflows that stick give something back quickly
As I was doing my post-mortem on my various failed AI habits, I considered the new behaviors that have stuck. Exhibit A: my compound writing plugin.
It’s a toolbox of skills I built for writing, modeled after Cora general manager Kieran Klaassen’s compound engineering plugin. It follows the steps of the writing process to help me take an article from idea to polished draft.
Compound writing had a few points against it, if we’re using my failures as precedent. I borrowed it from someone else, like Tend. It’s effort-intensive, like Margot. And it also has to compete with my established Slack-checking habit—as I’ve said I will do anything to avoid working on a draft.
And yet, compound writing has become the core of how I do my job.
The difference is that compound writing quickly started sending signals back that my effort would be worth it. Even before it worked well, the plugin made writing feel propulsive again. It gave me a paragraph to push against, a question that unstuck an argument, or a way into the next stage of a draft. It gave me a writing partner that solved a real problem in my work—the loneliness inherent to writing for a living.
In a series of studies, behavior researchers Kaitlin Woolley and Ayelet Fishbach found that immediate rewards predicted persistence better than delayed rewards. The research is not proof that a behavior will become automatic, but it helps explain why I kept at it. Compound writing paid me in currency that mattered to me immediately. The distant promise of becoming a calmer Slack-checker could not compete.
Solving AI with more AI (again)
Looking across my AI workflows—the ones that failed and the ones that stuck—I can see why each ended up where it did. Attention Desk solved a problem I mostly imagined for a person I hoped to become. Compound writing met me inside work I already did and began paying me back almost immediately.
I could judge a workflow by what happened once it made contact with my day-to-day work—whether I resented maintaining it, carried away something useful from its failure, or kept returning to it. But the worry that I’m missing out on something hasn’t gone away. I wanted a way to know for sure if an experimental system would never work, or if it could be useful with a tweak to its design or implementation.
So, naturally, I added a new AI system to audit my AI systems. This time, a reviewer called Agent Ops.
Agent Ops is a plugin inside my Codex that inventories the systems I built or adopted and asks me questions about the goal, usage, and the outcome of each one before making a recommendation to either keep, tweak, or retire it.
In setting up the plugin, I discovered that my failed workflows fall into four buckets: I forgot it existed, it failed to run, the output required too much maintenance, or it didn’t solve a recurring problem.
I turned those answers into provisional rules. A new automation has to run manually three times before I schedule it. Its output has to produce something I use with less than five minutes of review. Four unused runs of an automation trigger a decision to change or retire it.
The prompt below is a version of the interview that produced it, adapted so you can run it on your own Island of Misfit Workflows.
Become a paid subscriber to Every to unlock this piece and get Katie’s prompt. It'll interview you about every AI tool you built and stopped using, and tell you what to do with each one: Keep it, redesign it, wait for the right moment, or let it go.












