Reproducibility

Making Work a Transferable Learning Practice

Reproducibility is usually presented as a property of research: another person should be able to understand how a result was produced and, given the necessary materials and procedures, regenerate it. In teaching, I use reproducibility for a broader purpose. A quantitative result has a history involving evidence, transformations, analytic choices, interpretation, and revision, and students learn more when that history remains available for inspection. Reproducible work gives them a concrete way to trace how a claim came to be supported and where uncertainty or error entered the process (Sandve et al., 2013; Wilson et al., 2017).

My goal is for students to treat reproducibility as a practice of making intellectual work legible. They should be able to identify where information came from, explain consequential decisions, connect results to claims, revise work without losing its history, and leave a project in a condition that another person can understand and continue. Reviews of open and reproducible scholarship in education describe related benefits for scientific literacy, collaboration, writing, problem solving, and statistics instruction while also emphasizing that the evidence base remains uneven across outcomes and study designs (Pownall et al., 2023; Dogucu, 2025). I therefore treat reproducibility as an instructional environment in which these forms of reasoning can be practiced and observed.

Making the Process of Knowing Visible

A finished result can conceal the decisions that made it possible. A reported mean depends on how the sample and outcome were defined, a composite score depends on coding and missing-data decisions, and a regression coefficient depends on variable construction, model specification, and assumptions. Reproducible practice preserves enough of this chain to make the result inspectable as the product of a sequence of choices. Students can then see statistical output as one stage in an argument that begins with evidence and ends with a claim (Sandve et al., 2013; Wilson et al., 2017).

That visibility changes how I teach provenance. When students ask where a number came from, I want the question to lead toward how the underlying evidence was produced, transformed, and interpreted. The source of a result matters because the reliability of the pathway constrains how much confidence the result deserves. Research on epistemic cognition similarly emphasizes students’ understanding of how knowledge is developed and justified, with stronger epistemic cognition associated with conceptual understanding and argumentation (Greene et al., 2018).

A reproducible workflow gives those questions an object. Students can move backward from a conclusion to the output that supports it, from the output to the analysis that generated it, and from the analysis to the data and decisions on which it depends. They can ask whether a variable represents the construct they think it represents, whether an assumption is plausible, whether an exclusion changed the result, and whether the language of the conclusion fits the design and uncertainty. The quality of the final claim becomes connected to the quality of the reasoning that produced it.

Using the Visible Process for Feedback and Revision

Reproducible work also functions as external memory. Complex projects quickly exceed what students or researchers can reliably retain about file versions, coding decisions, abandoned approaches, model changes, and the sequence of revisions. Documentation, scripts, project structure, and version history allow consequential information to persist outside memory, which reduces the need to reconstruct a project from recollection alone (Markowetz, 2015; Risko & Gilbert, 2016). This makes it easier for students to return to earlier work and explain how the current version developed.

The preserved process gives feedback a more precise target. Feedback can address the task, the process used to complete it, and the learner’s regulation of that process (Hattie & Timperley, 2007; Butler & Winne, 1995). When the workflow remains visible, I can distinguish an error in data preparation from an error in analysis or an overstatement in interpretation. Students can then locate the point where the reasoning diverged and revise the work at that point.

Revision becomes part of the intellectual record. A changed coding rule should propagate to the analysis and interpretation, and a changed interpretation should remain accountable to the evidence that supports it. Students can compare what they expected with what occurred, examine why a result changed, and identify which earlier decision needs reconsideration. Over time, these records give students evidence about their own habits of reasoning and provide me with evidence about whether feedback from one problem is influencing how they approach the next.

I want students to become increasingly able to diagnose their own work. That includes recognizing when a result cannot be regenerated, when a claim exceeds the evidence, when a manual step has created uncertainty, and when documentation is too weak for another person to follow. The workflow becomes a surface for self-monitoring because students can compare their understanding with an inspectable record of what they actually did. This makes self-correction more concrete and gives revision a clear relationship to evidence.

From Following a Workflow to Owning It

Students can follow a well-designed reproducible workflow without understanding why it works. Early structure is still useful because expert reasoning is often invisible to novices. Cognitive apprenticeship emphasizes making expert processes observable through modeling, coaching, scaffolding, articulation, reflection, and increasing opportunities for independent action (Collins et al., 1989). In quantitative instruction, this means making visible why I question a variable, inspect an assumption, reject an analytic option, or qualify a conclusion.

I use templates and structured workflows to expose consequential decisions while students are still learning how to manage them. Early in a course, I may specify file organization, documentation conventions, or the sequence of an analysis so that students can devote attention to the reasoning those structures support. As their capability develops, I can transfer more responsibility for choosing, adapting, and defending the structure itself. A student who initially follows a supplied workflow should eventually be able to explain which parts are necessary, which parts are contextual choices, and what would need to change in a new project.

Ownership becomes visible when students can justify and modify the process. They should be able to explain why a transformation was performed, how an alternative would affect the analysis, what information another researcher would need, and how a project should be reorganized when its purpose changes. Reproducibility then supports research judgment because the student is responsible for a chain of decisions whose consequences remain visible. The aim is a workflow the student understands well enough to critique, repair, and reconstruct.

Generative AI makes this issue especially visible because polished text, code, and analysis can be produced with limited evidence of the learner’s underlying understanding. The instructional response I find useful is to preserve provenance and responsibility: students should know what was generated, what was verified, what evidence supports the result, which decisions they made, and where uncertainty remains. Current higher-education research on AI similarly emphasizes verification, critical thinking, and the risks of uncritical cognitive outsourcing alongside potential instructional benefits (Salido et al., 2025).

Teaching for Transfer

The broader value of reproducibility depends on transfer. A student who learns one workflow in one quantitative course may reproduce that procedure accurately without recognizing the same underlying problem in a different setting. Transfer varies with differences in domain, context, function, and required performance, so similarities that are obvious to an instructor may remain hidden to a learner (Barnett & Ceci, 2002). I make the transferable principles explicit and give students opportunities to recognize them in different forms.

Comparison is one way to make those principles visible. Research on case comparison shows that examining structurally similar examples can help learners abstract relationships that are harder to see from a single case (Alfieri et al., 2013). I can ask students to compare provenance in data analysis with source tracking in scholarly writing, or version history in code with revision history in a manuscript. Across those cases, students need to preserve origins, understand transformations, connect evidence to claims, and explain consequential decisions.

Writing provides a particularly close connection. Reproducible empirical work and scholarly writing both develop through cycles of evidence, interpretation, revision, and communication, and process-oriented instruction can integrate documentation, analysis, and writing as parts of one research practice (Marshall & Underwood, 2019). Similar principles appear in collaboration and professional work, where another person may need to understand why a decision was made, continue a project, evaluate uncertainty, or revise a conclusion when new evidence appears. Students benefit when these connections are named and practiced across contexts.

My standard is that students leave with questions they can carry into unfamiliar work. Where did this information come from? What happened to it? Why was this decision made? What evidence supports the claim? What changed, and what depends on that change? What remains uncertain? Could another person understand and continue the work? When students can use those questions to construct and evaluate their own processes, reproducibility has become part of how they reason across courses and professional settings.

References

Alfieri, L., Nokes-Malach, T. J., & Schunn, C. D. (2013). Learning through case comparisons: A meta-analytic review. Educational Psychologist, 48(2), 87–113. https://doi.org/10.1080/00461520.2013.775712

Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637. https://doi.org/10.1037/0033-2909.128.4.612

Butler, D. L., & Winne, P. H. (1995). Feedback and self-regulated learning: A theoretical synthesis. Review of Educational Research, 65(3), 245–281. https://doi.org/10.3102/00346543065003245

Collins, A., Brown, J. S., & Newman, S. E. (1989). Cognitive apprenticeship: Teaching the crafts of reading, writing, and mathematics. In L. B. Resnick (Ed.), Knowing, learning, and instruction: Essays in honor of Robert Glaser (pp. 453–494). Lawrence Erlbaum Associates. https://doi.org/10.4324/9781315044408-14

Dogucu, M. (2025). Reproducibility in the classroom. Annual Review of Statistics and Its Application, 12, 89–105. https://doi.org/10.1146/annurev-statistics-112723-034436

Greene, J. A., Cartiff, B. M., & Duke, R. F. (2018). A meta-analytic review of the relationship between epistemic cognition and academic achievement. Journal of Educational Psychology, 110(8), 1084–1111. https://doi.org/10.1037/edu0000263

Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. https://doi.org/10.3102/003465430298487

Markowetz, F. (2015). Five selfish reasons to work reproducibly. Genome Biology, 16, 274. https://doi.org/10.1186/s13059-015-0850-7

Marshall, E. C., & Underwood, A. (2019). Writing in the discipline and reproducible methods: A process-oriented approach to teaching empirical undergraduate economics research. The Journal of Economic Education, 50(1), 17–32. https://doi.org/10.1080/00220485.2018.1551100

Pownall, M., Azevedo, F., König, L. M., Slack, H. R., Evans, T. R., Flack, Z., Grinschgl, S., Elsherif, M. M., Gilligan-Lee, K. A., de Oliveira, C. M. F., Gjoneska, B., Kalandadze, T., Button, K., Ashcroft-Jones, S., Terry, J., Albayrak-Aydemir, N., Děchtěrenko, F., Alzahawi, S., Baker, B. J., . . . FORRT. (2023). Teaching open and reproducible scholarship: A critical review of the evidence base for current pedagogical methods and their outcomes. Royal Society Open Science, 10(5), 221255. https://doi.org/10.1098/rsos.221255

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002

Salido, A., Syarif, I., Sitepu, M. S., Suparjan, Wana, P. R., Taufika, R., & Melisa, R. (2025). Integrating critical thinking and artificial intelligence in higher education: A bibliometric and systematic review of skills and strategies. Social Sciences & Humanities Open, 12, 101924. https://doi.org/10.1016/j.ssaho.2025.101924

Sandve, G. K., Nekrutenko, A., Taylor, J., & Hovig, E. (2013). Ten simple rules for reproducible computational research. PLOS Computational Biology, 9(10), e1003285. https://doi.org/10.1371/journal.pcbi.1003285

Wilson, G., Bryan, J., Cranston, K., Kitzes, J., Nederbragt, L., & Teal, T. K. (2017). Good enough practices in scientific computing. PLOS Computational Biology, 13(6), e1005510. https://doi.org/10.1371/journal.pcbi.1005510