As a result, bringing data together into a unified format often requires significant integration effort.
Case Study
Coffee Roast Pilot (Analogy for Batch Processes)
Coffee roasting and brewing as an analogy to biologics batch processes to demonstrate multivariate analysis with AI. The goal was to understand what method led to the best brewed coffee. The coffee was characterized for aroma, taste, flavor, bitterness, acidity, and everything that described the bean themselves.
Three data types:
1. Process Variables
- Roaster temperature
- Drum Speed
- Time
2. Analytical measures
- Refractometer,
- pH
3. Metadata
- Grind size
- Water temperature
- Perceived flavor
After 10 runs, the pilot revealed that insufficient batch count led to underpowered statistical conclusions and required expanding the dataset. After 20 runs a model was able to achieve statistical relevance, which shows the importance of getting adequate data in order to have a successful model.
Automate Data Capture at the Source
“NIH studies show that 3 to 9% of manually transcribed data is erroneous.”3-5
Automate data capture at the source whenever possible to avoid transcription errors and duplicate records. Many legacy capture methods, including manual copy and paste, flash drives, and serial terminals, introduce risk and data loss.
Use standardized drivers and connectors that support automated uploads and common communication protocols; this reduces duplication and simplifies onboarding of heterogeneous instruments.
Retain Failed Runs and Out-of-Specification Data
Do not discard failed runs or out-of-specification results. AI models trained only on successful experiments cannot learn to recognize failure modes or process deviations.
Including the full range of Design of Experiments (DOE) conditions, including failed experiments, helps create more representative models and reduces the risk of biased conclusions. If data must be excluded, such as in cases of instrument malfunction, document the reason and retain that information as metadata.
Retain negative and out-of-spec data. AI models trained only on successful experiments will not learn to recognize failure modes and may hallucinate plausible but incorrect values.
Including full range of Design of Experiment (DOE) conditions, including failed experiments, helps avoid biased inferences. Make sure to annotate reasons for exclusion if any data must be removed (e.g., instrument malfunction) and store the annotation as metadata.
Ensure Data Provenance, Audit Trails, and GLP/GMP Alignment.
Capture key contextual information alongside the data, including timestamps, instrument unique identifiers, method versions, operator information, and transfer logs. This information, often referred to as “data provenance,” supports reproducibility and regulatory review.
As organizations adopt more digital workflows, regulators are placing greater emphasis on audit trails and compliance with GxP requirements. For critical datasets used in model development, organizations should also verify data integrity through mechanisms such as checksums, hashes, or signed logs.
Regulatory Considerations
Regulatory expectations in scientific, clinical, and manufacturing environments continue to evolve as organizations transition from paper-based workflows to fully digital ecosystems. Under the broader GxP framework, encompassing GLP, GCP, and GMP, regulators increasingly require organizations to demonstrate complete traceability, integrity, and provenance of all scientific data.
These categories span “C for clinical, M for manufacturing, and L for laboratory” and vary in strictness, with GLP generally less stringent than GCP and GMP. Nevertheless, across all domains, regulators now scrutinize not only the final data but also the processes, systems, and controls that generate and transport it.
Depending on theGxP or non-GxP requirements, data capture may require 21 CFR Part 11 compliance, which requires demonstrable traceability, immutability, and attribution. Audit files should be tamper-proof ensuring that even internal personnel cannot alter historical records. This design prevents tampering, supports trust in the data lifecycle, and aligns with regulatory requirements.
Collaborate Closely with IT and Cybersecurity
Engage IT early because they are a key part of your digital transformation project succeeding. Opening ports, setting firewall rules, and access to process networks require IT coordination.
Negotiate one-way data flows between segmented networks, where it is possible to reduce risk. IT departments are the gatekeepers and working with them is essential to avoid outages and security incidents.
Assure them and demonstrate that this process will be secure and will avoid insecure transfer methods (uncontrolled USB drives) and implement secure, auditable transfer mechanisms that have not capability of altering the data.
Pilot, Quantify, and Iterate
Run pilots to estimate sample size. In the coffee example, we found that it requires at least 20 replicates to reach statistical significance. Use pilot studies to determine the number of batches, replicates, and parameter ranges needed.
Measure data quality metrics (missingness, duplication rate, transcription error rate) and set thresholds before scaling.
Standardize Formats and Units
Normalize units and method identifiers at ingestion. Record whether values are measured by volume (typically done at lab scale) or by mass (typically done at larger scales) or adjusted and convert to canonical units before transfer.
Standardizing to mass for product additions ensures that formats are standard across scale. Adopt controlled vocabulary for method names, sample IDs, and site identifiers to avoid semantic mismatches.
Conclusion
People and process matter as much as technology. Successful data programs require domain experts, data engineers, and IT to collaborate and agree on priorities.
Budget for data engineering, instrument drivers, secure connectors, and validation are non-trivial investments but are essential to avoid downstream model failures. The regulatory landscape demands unified, tamper-proof, and fully traceable scientific data systems.
For models that will inform regulated decisions, design data capture with auditability and traceability from day one. AI models can deliver transformative insights into drug manufacturing, but only when fed complete, traceable, and representative datasets.
Teams should define use cases and success metrics first, automate and standardize data capture, preserve failed and boundary data, and work closely with IT to ensure secure, auditable pipelines. Iterative pilots will reveal the true data volume and quality requirements needed for robust models.
About the Author
Brendan Lucey, Chief Commercial Officer, Kivi Bio.
References
- Rand B. Bad data, bad results: when AI struggles to create staff schedules. Harvard Business School Working Knowledge. February 27, 2025. Accessed June 24, 2026. https://www.library.hbs.edu/working-knowledge/bad-data-bad-results-when-ai-struggles-to-create-staff-schedules
- Edjlali R. Lack of AI-ready data puts AI projects at risk. Gartner Newsroom. February 26, 2025. Accessed June 24, 2026. https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk
- Feng JE, Anoushiravani AA, Tesoriero PJ, Ani L, Meftah M, Schwarzkopf R, Leucht P. Transcription error rates in retrospective chart reviews. Orthopedics. 2020;43(5):e404–e408. doi:10.3928/01477447-20200619-10
- Mays JA, Mathias PC. Measuring the rate of manual transcription error in outpatient point-of-care testing. J Am Med Inform Assoc. 2019;26(3):269–272. doi:10.1093/jamia/ocy170
- Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. The FAIR guiding principles for scientific data management and stewardship. Sci Data. 2016;3:160018. doi:10.1038/sdata.2016.18