top of page
Search

Notes on Craft: Starting out with realist coding: A process of bucket sorting

  • rabrams0
  • 3 days ago
  • 3 min read

September 2026


In this latest blog Alankrita and I share how we started off coding in our review. We share the preliminary steps we took, and what helped us to work together on this. Our October blog will further ideas here and discuss in more detail how we formed our CMOcs, based on this coding approach.


Before starting analysis proper, it is a good idea to select two to three data points (i.e. three papers from the realist review searches), and code these together/ in a group. This helps ensure a degree of understanding across the data, and develop a loose coding framework. If you’re going in deductively initially, then you need to know whether this approach will suit the data, or if any changes need to be made. If going in inductively, you still need to build up an approach that will work across the whole dataset eventually. And, on top of this, you need both approaches to also draw on retroduction- because the point of a realist review is theory driven. You are seeking to move from your IPT to a final PT and so you need your realist logic of analysis to provide you with this knowledge.



Here are a few things Alankrita and I agreed upon when coding our initial three papers together:


  1. Code generously and into conceptual buckets. Large chunks, two to three paragraphs at a time are best, not line-by-line because this won’t provide any explanatory value when you come to revisit it for CMOc building.

  2. Avoid double coding. The same data wouldn’t typically sit in more than one code. It leads to lots of confusion so try to be pragmatic and cut the cake in one place. Where data overlaps across codes, make a decision and apply it consistently, as consistency matters more than perfect categorisation.

  3. Do not be tempted to start generating big buckets of contexts, mechanisms and outcomes. Separating these out in this way will not give you the explanatory value you need. Instead, start with broad descriptive labels, staying close to the paper's language. Watch out for labels that become too abstracted too quickly as these risk missing detail and may make it harder to form specific configurations later.

  4. In addition to data coding (which we did in NVivo), after reading each paper in full, write one to two configurations based on its key message. These will be grounded in what the paper actually says, not hypothetical. Write loosely at this stage; they will be refined iteratively. Store as a running bank in a Word document: author name or paper number, followed by the configuration statement. 

  5. When reading the next paper, ask: does this say something similar? Can I cluster this evidence to add more weight behind an existing statement? By about half way through your dataset, new papers should be adding evidence to existing configurations rather than generating new ones. Configurations need to be specific enough to say something concrete about the review at hand, but generalisable enough to apply across the dataset too.

  6. Realist coding is creative. There is a degree of abstraction needed so that evidence can be clustered into meaningful theories.

  7. Exact language match with the IPT is not needed, similarity of idea is enough. The IPT is a reference point, not a constraint.  

  8. The coding framework and configurations will all develop iteratively.


Alankrita's reflections:


In our double coding exercise, Ruth and I independently coded the same three papers and then met to compare and discuss our interpretations. I found this extremely valuable, as it enabled us to bring together our respective ideas and perspectives, develop a shared understanding of the emerging coding approach, and clarify how individual codes could be interpreted and applied.


For coding we used NVivo, which provided a practical way to organise and manage a large volume of data. Creating codes and sub-codes helped me structure the analysis, while annotations allowed me to record the rationale behind coding decisions. The ability to highlight coded text helped me avoid inadvertent double coding, and reviewing data within individual codes made it easier to assess whether the initial codes and sub-codes needed any refining. Overall, I found that these features supported a more organised and iterative approach to realist analysis and enabled us to be transparent and accountable in our methods.



Plus another we couldn't resist: Duddy and Wong, 2022.


If citing this blog, please do so as: Abrams, R and Singh, A (2026). Notes on Craft: Starting out with realist coding: A process of bucket sorting. September 2026.

 

 
 
 

Comments


Realist health & social care workforce SIG

  • alt.text.label.Twitter
  • alt.text.label.Twitter

©2023 by Realist health & social care workforce SIG. Proudly created with Wix.com

bottom of page