Last updated: 2026-09-30
Grounded Theory at AI Speed: Borrowing Agile Patterns for a Domain That Won't Hold Still
Grounded theory was built for domains that sit still long enough to be understood. A researcher gathers data, codes it, gathers more data to test the emerging categories against, and keeps cycling until new data stops producing new properties — the point Glaser and Strauss named theoretical saturation [1]. That cycle assumes the phenomenon under study changes slowly relative to the pace of the research itself. Study how experienced nurses triage patients, or how small businesses adopt a new accounting practice, and the underlying phenomenon will still resemble itself by the time the write-up is finished.
Study how people use large language models, and it won't. A codebook built around "prompt engineering" can be half-obsolete by the time a paper clears review, overtaken by test-time compute, reasoning models, and agentic tool-use pipelines that didn't exist when data collection started. Classic saturation assumes the target holds still while the researcher converges on it. In this domain, the target is also moving.like chasing a moving target
What Grounded Theory Assumes, and Where the Assumption Breaks FoundationalKnowledge that endures for decades — core principles
The method's core discipline is well established. Open coding breaks the data into concepts, properties, and dimensions with no prior category scheme imposed. Axial coding relates those concepts to each other — conditions, actions, interactions, consequences — building the categories into a structure. Selective coding integrates the structure around a core category, producing the theory itself [2]. Saturation is checked throughout by theoretical sampling: deliberately seeking out data likely to challenge or extend the current categories, and stopping only when it stops doing so [1].
Nothing about that discipline requires the underlying phenomenon to be stable — grounded theory was designed to build theory from data precisely because the phenomenon wasn't already well understood. What it does assume is that the phenomenon changes more slowly than the coding cycle converges. In AI and LLM research specifically, that assumption is the one that fails. A category built from careful axial coding of how users work around a model's context limit can be quietly invalidated by a context-window increase before the paper is drafted. The problem isn't that the domain is hard to study. It's that the ground the researcher is standing on to study it keeps shifting underfoot, at a pace the method's own convergence criterion was never built to outrun.cf. fractal terrain
the categories?"} D -->|yes| A D -->|no| E["Theoretical
saturation"] F["Domain shifts:
new capability,
new paradigm"] -.->|invalidates categories
before E is reached| B style E fill:#8FBF6A style F fill:#F2B8B5
Reconceiving Saturation as a Bounded Definition of Done Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Software teams facing a moving target long ago gave up on waiting for stable, complete requirements before shipping anything. Scrum's answer is the sprint: a fixed timebox, at the end of which a team delivers something meeting an agreed Definition of Done, rather than working indefinitely toward a moving notion of "finished" [4]. Applying that same discipline to grounded theory in a fast-moving domain means giving up on waiting for the categories to stop moving altogether, and instead asking a narrower, answerable question: has this sub-theory, bounded to this timebox and this slice of the phenomenon, stopped producing new properties from fresh data drawn from within that same window? That is local saturation — saturation of a bounded sub-theory or architectural behaviour, checked on a sprint cadence — as distinct from the global saturation classic grounded theory pursues across the whole phenomenon, which a genuinely fast-moving domain may not hold still long enough to ever reach.
This isn't as radical a departure from grounded theory's own practice as it first sounds. Guest, Bunce and Johnson's landmark study of interview-based saturation already operationalised the concept empirically, tracking how many new codes each successive interview produced and showing that the great majority of themes in their data set emerged within the first twelve interviews, with the rate of new-code discovery falling off sharply after that [3]. That is a measured diminishing-returns curve standing in for the judgement call "saturation has been reached." Treating the same curve as a live signal — the rate at which a coding pass turns up genuinely new properties, checked against a threshold at the end of each timeboxed increment — extends that precedent from a post-hoc justification into an operational stopping rule for a single sprint's local scope. Connecting Scrum's Definition of Done to grounded theory's saturation criterion this way is this page's own move, not a technique either Glaser and Strauss or Schwaber and Sutherland set out to build.diminishing returns: new insights slow down
Team Consensus as a Second Dynamic Threshold Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Saturation isn't the only classic grounded-theory discipline built on an assumption of a stable target. Inter-coder consensus — independent coders converging on the same labels for the same data, reconciled until agreement is high and stable across individual codes, properties, and edge cases [2] — has the identical vulnerability. Holding out for that kind of fine-grained agreement produces the same theoretical lag as holding out for fine-grained saturation: by the time two coders have reconciled every edge case, the phenomenon they're reconciling it about has moved on.
The fix is the same abstraction-ladder move applied to a different discipline. Rather than treating the required bar for agreement as fixed, scale it against how fast the domain is actually moving:
| Domain velocity | Where consensus is required | What counts as disagreement |
|---|---|---|
| Low (the classic case) | Micro-level: individual codes, properties, sub-categories, edge-case behaviour | Any unreconciled label difference blocks closure. |
| High (LLM and agentic-AI research) | Macro-level: the core relational structure and its primary interactions only | Micro-level variation is logged as temporal variance — a property of the category, not a reason to delay it. |
That reframes the question a team is actually answering at a sprint boundary. It stops being "do we agree on every label" and becomes "do we agree this high-level construct explains the phenomenon as it currently stands, and is it stable enough to publish as a baseline." Scaling the required consensus bar against domain velocity this way is this page's own extension of the sprint-bounded saturation idea above, not a documented technique in either the grounded-theory or Agile literature.
A scheduled team reconciliation point does double duty beyond simply settling disagreements faster. An individual researcher working alone is prone to getting caught up in whatever the newest model release or capability happens to be — a high-frequency signal that may or may not mean anything structurally. Routing observations through the whole team at a fixed sprint boundary, rather than letting consensus emerge whenever it happens to, filters that noise: a pattern only one researcher has noticed stays a candidate, and a pattern the team independently converges on becomes a category. When two coders do split on how to divide a category under time pressure, the tie-breaking rule that keeps a fast-moving research programme usable favours consolidation over proliferation — merging the split, not forking it — since runaway proliferation of ever-finer categories is exactly the endless-refactoring failure mode a velocity-aware framework exists to avoid.
observation"] --> G{"Sprint-boundary
consensus gate"} R2["Researcher B's
observation"] --> G G -->|independently
replicated| C["Stable macro
category"] G -->|single-researcher
only| N["Logged as candidate,
not yet a category"] style C fill:#8FBF6A style N fill:#FFF3CD
The practical stopping test that falls out of this is predictive rather than definitional: the team has reached working consensus once its members, categorising new incoming data independently, agree on most of it — not once every taxonomy boundary has been argued to a close. A concrete illustration of the trade-off: if a domain's working assumptions turn over roughly every three months, which is a plausible rough estimate for how fast agentic-AI tooling conventions have been shifting recently, a team spending six months chasing 95%-plus micro-level agreement on a fine-grained taxonomy publishes a theory about a phenomenon that has already moved on. A team settling for roughly 80% macro-level agreement inside a three-week sprint publishes something the field can still use while it's current, and treats the unresolved 20% as the seed of the next sprint's refactoring rather than a defect to clear before anything ships.
Mapping Agile Constructs onto Grounded-Theory Practice Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The mapping holds at more than one point, which is what makes it worth taking seriously rather than treating as a one-off analogy:
| Agile construct | Grounded-theory equivalent | What it does in a fast-moving domain |
|---|---|---|
| Backlog | The pool of uncoded data — transcripts, logs, benchmark runs, literature not yet reviewed | Makes explicit that not all evidence gets analysed at once, and lets the queue be reprioritised as new theoretical gaps open up. |
| Sprint / timebox | One bounded coding-and-analysis increment | Forces a decision point before data collection and refinement can run on indefinitely chasing a moving target. |
| Minimum viable theory | A published mid-range theory, scoped explicitly to the current paradigm | Gets findings into circulation while they're still accurate, rather than holding out for a completeness the domain won't sit still for. |
| Refactoring | Re-running axial or selective coding on existing categories | Updates higher-level categories when a new capability or paradigm invalidates the lower-level constructs they were built on. |
| Pivot | A formal, documented narrowing or shift of theoretical scope | Names the moment a foundational assumption breaks (a shift from prompt engineering to agentic workflows, say) as a deliberate scope decision, not a quiet abandonment of earlier work. |
uncoded data"] --> SP["Timeboxed
coding sprint"] SP --> V{"New-code rate
below threshold?"} V -->|no| SP V -->|yes| MVT["Publish minimum
viable theory"] MVT --> BL SP -.->|foundational assumption
breaks| PIV["Pivot:
redefine theory scope"] PIV --> BL style MVT fill:#8FBF6A style PIV fill:#FFC857
Separating What Decays Fast from What Doesn't FoundationalKnowledge that endures for decades — core principles
Not every category a fast-moving domain produces is equally fragile. A category built around a specific prompt phrasing, a specific model's quirks, or last quarter's state-of-the-art benchmark score decays quickly, because the artefact it describes is itself ephemeral. A category built around how people recover from a tool's failure, how trust shifts after a visible error, or how responsibility gets negotiated between a human and an automated system tends to persist across model generations, because it describes a pattern in the interaction rather than a property of one implementation. Deliberately sorting categories into these two layers while coding — and aiming the theory itself at the slower-moving one — keeps a mid-range theory useful past the lifespan of whichever model or technique happened to be current when the data was gathered.
Two further habits protect that separation in practice. Treating a code's timestamp, and the model generation it was collected under, as a property recorded within the coding matrix — rather than as an embarrassment to be smoothed over — turns an apparent contradiction between old and new data into a dimension of the category itself: the theory doesn't just say what happens, it says what happens under GPT-4-era tool use versus what happens once agentic pipelines are common. And publishing in increments — working papers and scoped mid-range theories released as each timebox closes — spreads the risk that any single release goes stale, rather than concentrating years of unpublished synthesis into one monolithic theory that a single paradigm shift can retire overnight.
Where the Analogy Breaks Down FoundationalKnowledge that endures for decades — core principles
This framing has real limits, and they matter more here than they would for a looser analogy.
- There is no product owner for a scientific theory. A sprint review has someone in the room with the authority to accept the increment as done. A grounded theory's real acceptance test is the wider research community — peer review, replication, citation by others working the same seam — and that feedback loop runs on a timescale of years, not two weeks. A category that survives a sprint-bounded local saturation check has cleared a much lower bar than one the field has actually stress-tested, and reporting the two with the same confidence would be a mistake this framing doesn't excuse.
- A velocity metric can be gamed, including by accident. Once "new-code rate below a threshold" becomes the criterion for calling a sprint done, a researcher under time pressure has an incentive to code more coarsely, or to stop looking as hard for disconfirming data — the same target-distorts-behaviour risk this site's discussion of Goodhart's law in rubric design raises for a different measurement. A velocity threshold is a useful trigger for a human's saturation judgement, not a replacement for it.
- Data quality doesn't reduce to queue depth. A software backlog's items are broadly comparable units of work. An interview transcript, a benchmark log, and a single line from a forum post are not comparable evidence, and treating the backlog metaphor too literally risks flattening genuinely different evidentiary weights into one undifferentiated queue.
Where This Is Already Being Tried Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
The broader idea of operationalising saturation, independent of any Agile framing, is an active and unsettled area rather than a solved one. A 2025 preprint proposes a machine-learning decision-support system that predicts an appropriate sample size in advance from ten input parameters — research scope, information power, and researcher competence among them — rather than monitoring code emergence during analysis itself [5]. That's a different mechanism from the sprint-bounded, in-progress velocity check described above — it operates before data collection rather than during it — but it's evidence that treating saturation as something other than pure researcher intuition is a live methodological question well beyond this page, not a niche concern invented for AI research specifically.
Related Topics
- Project Marking as BDD — the same move in a different direction: borrowing a software-engineering discipline (Given/When/Then) to sharpen judgement in a domain that isn't software.
- Evaluating Detailed Rubrics — the Goodhart's-law risk of a measurement becoming a target, raised there for marking criteria and here for a saturation velocity metric.
- How LLMs Reason, Self-Correct and Get Checked — background on the reasoning-model and test-time-compute shifts used above as examples of categories a codebook can be overtaken by.
- Synthesis by Humans and Machines — the wider question of how synthesis work divides between a human researcher and an automated system.
References
- Glaser, B. G., & Strauss, A. L. (1967). The Discovery of Grounded Theory: Strategies for Qualitative Research. Aldine.
- Strauss, A., & Corbin, J. (1990). Basics of Qualitative Research: Grounded Theory Procedures and Techniques. Sage Publications.
- Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), 59–82. https://doi.org/10.1177/1525822X05279903
- Schwaber, K., & Sutherland, J. (2020). The Scrum Guide. https://scrumguides.org/scrum-guide.html
- Tutar, H., Erden, C., & Şentürk, Ü. (2025). Q-Sat AI: Machine learning-based decision support for data saturation in qualitative studies. arXiv:2511.01935 (preprint).