Part 1 of 2
AI can generate faster. But can it generate better? In L&D, fluency is not the goal; significance is. This article explores why L&D teams need to move from passive review to active discernment in AI-enabled workflows, especially before fluent output becomes locked into design decisions. Read on to explore how systems design, not just prompt engineering, is the real solution for responsible, high-impact learning.
At some point in the last two years, something shifted. Not dramatically. Not all at once. But the shift compounded.
I have observed three changes with compounded effects.
Fluency and reliability used to go hand in hand. If you’d really done the work, it showed in how you spoke to it. And if you hadn’t, that showed too. You’d hedge. Or get defensive. Or go very general very quickly. Those weren’t just communication styles. They were signals. They told us where to press, what to question, when to trust
AI-generated output is fluent in design. Confidence is mostly a function of language, not necessarily evidence.
When producing a thorough analysis took time and effort, it warranted a second or even third pair of eyes. The effort made the review feel proportionate. Now that it’s instant, nobody decided to stop. It just became optional.
When a human analyst was wrong, you could ask how. You could trace it back. With AI output, you can spot that something is off. You can fix the answer. But you can’t fix the reasoning. Which means the same error can recur, and you won’t necessarily see it coming.
These three shifts together are what this is about.

“AI can make mistakes. Please double-check responses.”
Every LLM interface carries that statutory warning. We read it the same way we read the warning on a cigarette packet. We know, and we move on.
Of course we are aware of AI hallucinations. Source citations for research that don’t exist. Dates, facts, attributions, all confidently wrong. We’ve named it, and we’ve built workflows around it.
My problem is not about what AI gets wrong. It’s about what it’s built on.
We all know that LLMs were trained on enormous amounts of text from the web. But that text is not a neutral sample of human knowledge. It’s a sample of what got written, published, shared, and clicked. What was frequent, recent, and sounding confident. The model learns to surface that promptly, and with remarkable fluency.
What’s worth noting is that a significant chunk of that training data is commercial. Landing pages, product descriptions, and ad copy. Decades of copywriters using the same construction: “It’s not X. It’s Y.” Dismiss the ordinary reading, install the preferred one. LLMs absorbed that pattern because it was everywhere and it performed. It’s not a flaw. It’s the mechanism.
Psychologists call this the “processing fluency” effect. When information is easy to process — clear, well-structured, confidently expressed — we rate it as more truthful and more credible, independent of its actual accuracy. The experience of ease gets misattributed as a signal of truth. Alter and Oppenheimer documented this comprehensively. Schwarz showed the same effect operates specifically in persuasive and commercial contexts. It’s one of the most replicated findings in cognitive psychology.
LLMs didn’t create this tendency. They were trained on it, and they exploit it perfectly.

Fluency bias is a general problem with media and persuasion. What makes it specifically problematic for L&D, where practitioners are trained to think critically?
Here’s what I think.
Critical thinking in L&D is grounded in pedagogy, not content. We learn to interrogate a learning objective, push back on a stakeholder who confuses awareness with capability, question whether an activity produces the target behavior. We get good at detecting the specific kinds of flawed reasoning that appear in learning design.
What we are not trained to do is interrogate fluency itself.
Learning design gets reviewed, of course. But the problem is that the reviewer reads the same fluent output. Fluency doesn’t always get caught in review. Mostly, it gets confirmed.
How do you interrogate a learning design that sounds authoritative? I look for three things: Do the objectives name a behavior or just a topic? A topic tells you what the course is about. Behavior tells you what the learner will do differently on the job. Bloom’s verbs are particularly good at obscuring this distinction because terms like “analyze,” “evaluate,” “apply” can sound precise while pointing at nothing specific.
Can the criteria distinguish a strong response from a weak one? Good criteria create a gap between adequate and excellent that you can actually see. If your criteria would pass almost anyone, they’re not criteria. They’re reassurance.
Does the design decision connect to the identified gap, or does it cater to a preference? Every design involves choices such as activities, formats, and sequences. The question is whether those choices trace back to what the learner needs to do differently, or whether they reflect what the designer finds interesting, what the stakeholder requested, or what the platform makes easy. The difference is in reasoning, not the output.
If those hold up, the authority is grounded. If they don’t, you have fluency without substance. Consider two objectives for a technical course:
Grounded: The learner will diagnose a system error by working through a defined fault sequence before escalating.
Hollow: The learner will demonstrate proficiency in handling technical issues independently.
Both read as professionally written. The first describes a situation, a decision, and an action. The second sounds measurable but isn’t — proficiency in what, exactly? Under what conditions? The fluency is doing all the work that the thinking should be doing. And the cost of missing that is a program that likely produces confident learners who cannot perform.
There’s a distinction that cuts through this: salience is what stands out; significance is what matters. LLMs are salience engines, surfacing what was statistically prominent in training data. The job of an L&D practitioner is to tell those two things apart. And fluent AI output is best positioned to bypass exactly that.
This is not to say that practitioners aren’t thinking. We are. But we’re thinking about the content. The problem is in the mechanism.

There is a narrow window available to L&D practitioners to catch the pull of salience and anchor the design to the right gap. It is before the brief is frozen, before the structure emerges. Here’s a typical scenario.
A mid-sized organization. Sales numbers are down. The brief arrives before anyone has asked a question: the sales team needs to get better at communication. And they’re not wrong. Communication is a gap. But the term covers everything from email tone to difficult conversations to executive presence. Nobody stops to ask which part is actually driving the problem.
A learning designer gets the brief. It sounds coherent. The objectives write for themselves. Bloom’s verbs are applied. The course gets built, delivered, and rated well. But, the sales numbers don’t move.
What happened wasn’t negligence. “Communication” was the right territory. The significant gap that the team was pitching before confirming what the client actually needed was inside that territory and never made it into the design. That is something you can design for. A general communication course isn’t.
It required someone to ask what the learner would need to do differently, in which specific moment, and what was stopping them. Nobody asked. This is a classic case of salience mistaken for significance.
The brief was too fluent to create that pause. And in a workflow optimized for speed, agile development, rapid prototyping, and quick turnaround, there is only acceleration. What looked like a bottleneck before now looks like progress, and the window closes faster than it used to.
Prominent is usually right at the level of direction. Rarely is it sufficient at the level of design.

