In this essay, I will analyze the overlaps and discontinuities between my research data management practices for a study undertaken in my Human Information Interaction and what I have learned so far about the variously-described research data life cycle.
We’ve explored several variations of the research data management life cycle (RDMLC) this semester, each with their strengths and weaknesses. Corti et al. combine processing and analysis into a single step in their six-step model (Corti et al., 2020), while Thompson et al. combine dissemination and reuse into a single step in their six-step model (Thompson et al., 2023). For the purposes of this assignment, I will be using the non-combined elements of both:
- Plan
- Create and collect
- Process
- Analyze
- Disseminate
- Preserve
- Reuse
The case study
The research project is intended to explore the information experiences of plural systems and people with dissociative identity disorder (DID). Systems are characterized by having more than one independent identity sharing the same body. Despite stigmatizing portrayals in popular and news media, most systems are just ordinary people. Their information practices constitute an understudied and interesting intersection between that of individuals and that of groups. While much research has examined the experience of memory barriers in systems (inability of one identity to access the memory of another identity), it has exclusively focused on the neuro-psychological mechanisms and presentations, rather than centring the experiences of the participants. To my knowledge, no information behaviour research has examined the lived information experiences of systems.
For my study, my intention is to interview three systems, including as many individual identities within those systems as wish to participate. The semi-structured interviews will aim to elucidate the everyday information experiences of individual identities within the system, the system as a group, and the system presenting to the outside world as an individual. Of particular interest is the way systems curate their personally-relevant information and the way they handle exchanging information between inside and outside the system. Interviews will also include an Information World Mapping exercise. All interviews will take place online and be recorded for later coding and thematic analysis. As part of the consent process, participants are promised that their identities will remain completely private and will not be stored in the research data, and that all of their data will be destroyed when the research project is complete.
Overlap
I had the good fortune to already started taking a course on research data management when I started working on this research project. Returning to the 7-step process outlined above, the plan portion is well underway, with project decisions being documented in the project’s README file as they arise. File naming conventions and folder layouts already identified. I have created templates for the interview questions and transcripts along with a spreadsheet to hold their metadata, both based on the UK Data Service templates (UK Data Service, 2019). I am currently creating a spreadsheet to hold the codebook and themes for the analysis. The research proposal, consent form, and recruitment materials are also created.
The actual interviews will be starting after reading week, utilizing the prepared resources noted above.
Processing will consist of generating a codebook for the elements of the interview relevant for the research by reviewing the interview recordings. These will go in the codebook spreadsheet with references to which participants produced them and where in the recordings they are identified.
Analysis will involve taking the coded results and trying to identify themes where participants shared experiences in common and those where they diverged. Identified themes will go in the themes spreadsheet with references to which codes they are based on.
With an eye towards dissemination, preservation and reuse, I am utilizing open standard file formats wherever possible. Interview audio recordings will be created in Ogg Vorbis format, text documents will be created in ODT format, and spreadsheets will be created in ODS format. The presentation spreadsheet will be created in ODP format. Dissemination will take the form of a presentation to the class and a written report in the style of an academic journal research article. The publication requirements specify submission file formats that are not as open, necessitating conversion with retention of the originals. This will be accomplished by converting the ODT file of the report to DOCX, and the ODP of the presentation converted to PPTX and PDF. There is currently no plan to disseminate the research data beyond their use in producing the course assignment deliverables. Interview recordings and transcripts, codebook, and theme spreadsheet will not be disseminated, though they will be referenced in the presentation and report.
Gaps in the 7-step RDMLC
All data collected and processed for this research is stored securely in an encrypted folder on my password-protected computer in my locked apartment. Encrypted backups are created to an external hard drive, as well as an off-site location. Decryption of the data requires both a password and an decryption key, neither of which are stored with the data. This decision is not made lightly.
Given the small sample size and the intimate details shared during the interview process, sharing the interview recordings and transcripts cannot be done while maintaining the privacy of participants. Given the significant stigma around experiences of plurality, participant safety and anonymity is paramount. For that reason, the study is designed with the intention of destroying the research materials after the study is completed. At the conclusion of the research project, all data including all of its backups will be destroyed. Only the project documentation, templates, and the final report and presentation will be retained. This decision balances the needs of future researchers with the well-being of the participants.
Gaps in my own approach
Broader dissemination of research data is one of the goals of the research data life cycle. However, given the data destruction requirements of this study, there will not be any data to disseminate. There will also be no data that requires long-term preservation, though care is taken to preserve the data adequately during the research project. Similarly, there will be no opportunity for data reuse for secondary analysis, follow up research, or other applications.
Life Cycle Model Feedback
Neither of the research data life cycles considered here cover the destruction of data. While preservation and reuse of data for research is desirable in general, not every case is amenable to those outcomes. Especially in cases like this, with non-anonymizable research data from a very small population, a data destruction guarantee is sometimes a requirement to successfully recruit participants in the first place, much less satisfy ethic oversight requirements. While much attention is paid to the preservation and reuse of data, and rightly so, it seems that many research data management authors pay little attention to the rare but important case when destruction of that data is required. The secure destruction of data, including backups, is a potential concern even in cases where data is intended to be preserved and disseminated, for example when a study collects data incorrectly and must start again with a different questionnaire, the original invalid responses may need to be securely deleted.
Similarly, little attention seems to be paid to discarding research data once it is no longer useful; while it can be difficult to predict what future researchers may make use of, the needs of future researchers have to be balanced against the ongoing cost of curating, storing, and ensuring the integrity of research data. Not all data needs to be preserved forever.
Corti et al. choose to combine processing and analysis into a single step, but my experience thus far suggests they are better separated. Processing the data involves a variety of data cleaning steps that Thompson et al. spend two chapters covering. Conversely, Thompson et al. combine dissemination and reuse into a single step in their model, but again I feel these are better as separate steps. As a researcher, I might choose to continue my line of inquiry building on past data (reuse) but the consent for that data collection may not have included it being seen by anyone other than the research team (dissemination). Dissemination focuses on topics like discoverability, giving the dataset a good description and keywords, making sure it has a persistent identifier; reuse focuses on metadata documentation, open file format standards, and clear variable names. While there is significant overlap with final intention, i.e. having the data be discoverable and usable by future researchers, the direction of attention for each are dissimilar enough in my view that they should be considered separately.
Applications
This assignment has given me an opportunity to critically reflect on the needs of research participants within the research data life cycle, and how the assumptions we make in our efforts to support researchers can have wider impacts than might be readily apparent. While models like the one from the UK’s Data Curation Centre offer more complexity and depth, they are more difficult for researchers to understand and adopt (Higgins, 2008). I hope to include this critical reflection aspect when advising researchers about their data management plans, helping them centre the experiences and needs of participants, and understand edge cases that need to be planned for.
References
Corti, L., Eynden, V. van den, Bishop, L., Woollard, M., Haaker, M., & Summers, S. (2020). Managing and sharing research data: A guide to good practice (2nd edition.). SAGE.
Higgins, S. (2008). The DCC Curation Lifecycle Model. International Journal of Digital Curation, 3(1), 134–140. https://doi.org/10.2218/ijdc.v3i1.48
Thompson, E. by K., Hill, E., Carlisle-Johnston, E., Dennie, D., & Fortin, É. (Eds.). (2023). Research Data Management in the Canadian Context. Western University, Western Libraries. https://doi.org/10.5206/ZRUV7849
UK Data Service. (2019). UK Data Archive Model transcription template. University of Essex.