Coding Consistency: Why Re-Generating Categories Can Produce a Different Set Each Time

Qualitative Analysis
Reference
Updated Sep 25, 2026

Coding Consistency: Why Re-Generating Categories Can Produce a Different Set Each Time

If you've run AI Auto-Generate Categories twice on the same question and noticed the resulting categories aren't identical - different names, a different split of themes, maybe one run surfaced something the other didn't - that's not a malfunction. It's a real property worth understanding, especially if you're running a study more than once.

Why This Happens

Each generation run analyzes the responses available at that moment and drafts a fresh category set from them - see Running Multiple Classification Passes on the Same Question for the mechanics. A second run isn't guaranteed to reproduce the exact same category boundaries as the first, even against the same response set, because there's often more than one reasonable way to group similar responses into named themes.

When This Doesn't Matter

For a one-off study where you're only ever going to generate categories once, this isn't something you'll notice or need to think about - it only becomes relevant when you compare results across separate generation runs.

When This Matters a Lot

For a tracking study - the same question asked wave after wave, where you want to compare results over time - independently regenerating categories each wave risks producing category sets that don't line up cleanly, which makes wave-over-wave comparison unreliable even though each individual wave's results are perfectly valid on their own. See Comparing Category Sets Across Two Survey Waves for the practical fix: build your categories once, then reuse the same definitions manually for later waves instead of regenerating independently each time.

This Applies Within One Question Too

Even outside a tracking study, if two team members each independently generate categories for the same question - maybe to compare notes, or because one person's first attempt got discarded and someone else tried again - expect the two sets to differ somewhat rather than assuming one of them made an error.

What Stays Consistent

Classification itself, once run against a fixed, already-approved category set, applies that set's definitions consistently across every response in that run - the variability described here is specifically about generating categories fresh, not about how consistently an existing category set gets applied during classification.

FAQ

Is there a way to make two generation runs produce identical results?
Not directly - if you need identical categories across two points in time, the reliable approach is defining the category set once and reusing it manually rather than regenerating.

Does this mean AI-generated categories are unreliable?
No - each individual run produces a reasonable, defensible category set based on the responses it analyzed. The variability only becomes a problem when you need two independent runs to match each other exactly, which is a specific use case (like tracking studies), not a general reliability issue.

Should I be concerned if two people on my team get different categories generating separately?
Not concerned, just aware - it's worth agreeing on one category set as the team's standard (see Editing and Refining Categories in Text Analytics) rather than treating either independently generated set as automatically correct.

Does this variability affect classification confidence scores?
Confidence scores reflect how well a response fits the category set it was classified against - they're not a measure of how similar that category set is to a different, independently generated one.

For the practical steps to keep a tracking study consistent despite this, see Comparing Category Sets Across Two Survey Waves.

text analytics category generation coding consistency tracking study

We value your privacy

We use cookies and similar technologies to improve your experience, analyze site traffic, and personalize content. Learn more