← Writing

Writing · 10 min read

The AI 75% Trap

The 75% keeps getting easier. The 25% is where the work lives.

Russ Unger · May 11, 2026

A mountain's final summit ridge catching golden light, the hardest last stretch of the climb

As part of my Advisory role for Rockford University's Generative AI for Value Creation Certificate Program, I get asked to provide feedback on prompts to help improve the program. This article stems from some of those prompts, updated to fit this context better.

This is one of those "lessons learned" articles, however, it's not one of those where I imply or insinuate that the way I am working is the same way that everyone else works. I know what I'm learning and what patterns I'm seeing about how I work, and how the team I work with operates. If this resonates with you, that's great! And if it's contrary to how you work, I hope you'll share so that I can continue to get better and identify better paths, too.

For me, 0 to 75% is the easy part of generative AI. Any one of us can do that with speed and ease, and for me, it feels like the "hello world" of our era. Yeah, it's empowering and addictive, and sometimes it means there's a glowing screen keeping me awake late into the night to get just one more fix in.

The problem is that at 75%, people tend to feel this is an okay place to stop in their effort, and it isn't. I spend 25% of my time getting from 0 to 75% of a solution, and 75% of my time traversing that last 25% to completion of any AI effort.

That last 25% is where expertise, iteration, and judgment do their real work. And it's also where a lot of teams quit, or think that they've landed on a solution that is good enough. And yes, I'm stating this based upon any number of those amazing claims you see in Reels, in micro-content, and heaven forbid, Reddit posts. We're all still looking for that one-shot fix to financial freedom empowered by our own personal Open Claw with a credit card that does business while we sleep.

The gap between what AI makes look easy and what shipping AI-powered work at scale actually requires is the thing we should be talking about.

Guardrails & What They Might Look Like

The examples that hold up at scale for me share one feature: I made the deliberate choice to push past the 75% mark, and built guardrails to force that to happen. The most instructive cases are tools where AI augments my expert judgment rather than replaces it.

The clearest version of this in my own work has been with heuristics. The fast path is to grab Nielsen's ten, drop them into a skill or a prompt, and call it a day. I've not gone that route, and I'm a bit mortified by the assumption behind it, which is that ten general-purpose heuristics need no additional work to be valuable in any specific context. I'm sure that those skills will produce output that sounds confident and feels like rigor. That is, after all, what AI does very well: sounds confident and feels like rigor.

The problem is that ten general-purpose heuristics applied to any interface, any flow, any product context tend to produce ten general-purpose observations. Error messages should be helpful, well, no kidding, and what does helpful look like? Navigation should be clear, and clear without context is again not very helpful. While these are good overarching heuristics, they're not actually evaluating the thing in front of me.

I've found that it's not too terribly difficult to create sub-heuristics that can map categorically to the Jakob 10, or even to Morville's Honeycomb, or to the works of others, which can really help with the focus and granularity. It also requires more setup than I wanted to do initially; mapping heuristics into functional categories and subcategories, building context about the product type, the user, and the decision an interface is asking someone to make, and treating the prompt like a curation problem rather than a search query. That's the 75% to 100% work, and it's mostly invisible from the outside.

The same pattern shows up in the content guardrails I've built for tools that generate content with human editors. There are topic boundaries, tone calibration, what will and won't get published, and how to handle the moments where AI assistance is helpful versus where it would actively hurt the work. Those guardrails exist because I trust AI to do exactly what I tell it to do, including the wrong things if I haven't shaped the frame carefully.

The guardrail is the expertise itself, embedded in the system before the AI ever runs. The curation has had to happen up front every time. When I've skipped that step, the quality degrades the moment I get into testing. If I cannot get it to a point where I don't think any of you would make fun of me, I walk back through my process and do the work I tried to shirk off in the first place.

Or, even worse, I get fooled enough to try and ship something at scale, and get pretty terrifically humbled somewhere around the 10th or more prompt revision attempt, when I start to realize that the only way to fix the issue is to rewind the effort and get the quality guards in earlier in the process.

Ask me how I know.

What UX Professionals Need to Understand About GenAI & What GenAI Leaders Need to Learn from UX

I hate click-bait titles and sections headers like this, and yet...

What I've learned, watching myself and the team I work with operate in this space: AI does a remarkable job of fooling the ill-informed that it's doing a great job. On topics where I'm shallow, the output looks great to me, and I have to fight the urge to ship it. On topics where I'm deep, I can see what's off the moment I read it, and I know how to tune the prompting and the evaluation to get somewhere better. Domain expertise is the calibration mechanism that I must rely upon, and without it, I don't actually know what I'm shipping. That has made me a believer in leaning harder into the depth of my own discipline.

And it's made me an even bigger believer in the importance of Discovery, and in pairing the people who can turn concepts into outcomes with the people who know what those outcomes should be.

The lesson going the other direction is one I keep running into when I'm working on delivery rather than evaluation. Delivery at the scale of 1 is very different from delivery at the scale of 10. This is the big one, and my theory is that a lot of people kind of quit, or at least stop here longer than they should, because they haven't considered what happens when that single output gets repeated and repeated and repeated. And then somewhere along the way, that third, fourth, or fifth output starts to get glossed-over, then eventually ignored completely.

What sounds good for one piece of content starts to sound formulaic, basic, and without real intelligence the moment scale moves to multiple. I iterate to reach a steady state of 1, and then I iterate again, repeatedly, to reach a steady but variable state across many. UX has been thinking about consistency-without-sameness for decades. AI-first work that skips that lesson ships slop at scale, and I've shipped my share and fought with the machine to figure out what I was missing.

I keep coming back to the same conclusion: neither side gets to skip the other's discipline. UX work without AI fluency under-leverages the moment. AI work without UX rigor mistakes output volume for value. I sit in both rooms, and most of what I'm trying to figure out lives right in the space where they overlap.

AI Slop & the Trust Environment

The tension is real, and I'd reframe it. The deeper issue I keep running into is that the volume of AI output, where calling it mediocre or AI slop feels generous already, is actively eroding the trust environment for everyone, myself included. I've said it before, and I'll say it again: prompting without context delivers the mediocrity of the internet.

What I'm watching happen is user skepticism climbing rapidly, and the cause is straightforward. People are encountering more and more content that looks competent on the surface and doesn't hold up on a second read. It looks competent enough that any random person can feel confident dumping a wall of text into a forum online, right before getting rightfully lit-up by the resident experts who are already tired of getting their expertise AI-splained to them.

The lagging measurable is that the upper quality ceiling is coming down. If I put in half effort with AI, I get terrible results. If I put in maximum effort, I might maintain what audiences have been accustomed to. That's an asymmetry I haven't seen enough people factor into how they think about AI-powered work, and it's one I've had to factor into how I budget my own time. It also, frankly, kind of sucks to know that because others are taking shortcuts the rest of us have to work harder to maintain what might feel like a reasonable standard.

The principle that has held up for me is something I'm quoting my supervisor on, and it's incredibly simple: verify, then trust. AI doesn't have all the solutions, and it won't fill in the gaps out of the kindness of its heart (it's why your billion dollar idea didn't magically Open Claw itself into your wallet just yet). It won't identify the strategy I should pursue, and it won't see the blindspots I'm overlooking, at least not without a lot of poking, prodding, and asking on my part. It produces plausible output and lets me decide whether to ship it.

My best work in this space has come when I've treated AI output as a starting point that has to clear an expert bar before it goes anywhere. My worst work has come when I've assumed AI cleared that bar by default. It doesn't, and I've found that the biggest mistakes happen when I've been too comfortable, lazy, or both.

Investing in the Final 25%

Four investments matter most for getting past the 75% trap, and I notice them being underweighted in many of the conversations I'm in, or sometimes just having with myself. AI makes it seem like there's a shortcut to be taken for everything we do, and I understand the hype. I also understand that this is a longevity game, and the temporary players won't have the stomach for it.

  1. Domain expertise in the loop, not adjacent to it. The calibration mechanism I keep talking about only works when real subject matter experts are actually looking at AI output and saying "this is right" or "this is off." That's a fundamentally different job from data scientists evaluating model performance, and the two don't substitute for each other. When I've had domain experts inside the workflow (admittedly, many times this alleged expert is me), AI work has held up; when I've relied on technical evaluation alone, or assumed AI was checking itself, the results haven't been worth shipping.
  2. Budget for the last 25%. Most AI initiatives I've seen, including some of my own, are funded as if 75% complete is shippable. They generally aren't, and it's challenging because this is an exciting time and getting something right is less exciting than getting something MVP'd for the first time. That dopamine hit you get the first time your "Hello World" happens is addictive (and your friends are already tired of your neat little incomplete idea and don't want to see it every step of the way from there, either). The plan needs to include the iteration cycles to get from 75% to actual delivery, governance of the content over time, and the discipline to resist the cultural pull to call it done early.
  3. Scale testing as a discipline. If a solution works for one output or outcome, that tells me very little about how it works for 10 or 100. AI is amazing in that, when it sounds and acts just like us, we want every output to be uniquely the same as the last. I've learned to build in evaluation at the scale I intend to ship, far beyond the scale I can demo. The scale-of-1 demos are easy; they're also where I've fooled myself most often.
  4. "Verify, then trust" as a governance principle. I quoted this phrase from my supervisor earlier in the article, and at the org level it has to be a real workflow requirement to mean anything at all. The question I want every AI-powered touchpoint to answer is: who verified this, against what, before it reached an audience? When the answer is "no one, by default," I know I've got a problem upstream of the touchpoint.

The organizations that get this right will pull away from the pack. The ones that mistake speed for progress will spend the next few years watching trust erode under their own output.

The 75% keeps getting easier. The 25% is where the work lives.