What Direct Data Capture Actually Changes

Start with what does not change. Direct data capture is not a new category of software. The European Medicines Agency addressed this directly in its 2019 qualification opinion on eSource direct data capture, noting that an eSource system can be considered an EDC system and that EDC systems already allow for direct data entry when this is defined and approved in the trial protocol.2 The regulatory foundation is older still. The FDA published its guidance on electronic source data in clinical investigations in September 2013, and that document already anticipated data entered directly into the eCRF with no prior paper or electronic record.1

So the vendor pitch that direct data capture represents a technological leap is overstated. What is understated is the process consequence. A trial that declares certain eCRF fields to be the source record has changed its record model, and the record model is what every downstream clinical operations and data management process is built on.

The transcription model and the assumptions inside it

In the conventional model, a study coordinator observes a blood pressure, records it on a worksheet or in the electronic medical record, and later enters it into the eCRF. The worksheet or medical record entry is the source. The eCRF entry is a report of the source. Two records exist for the same observation, and they are created by different acts at different times.

This structure carries three assumptions that most clinical teams have stopped noticing. First, an independent copy of the observation exists outside the sponsor’s control, so an error introduced during transfer to the sponsor can be detected by comparison. Second, a monitor can determine whether the reported value is right by looking at a document. Third, a query about a value can be resolved by consulting that same document. The whole apparatus of source data verification, query resolution against source, and the standard data management plan is built on those three assumptions.

The direct capture model

In direct capture, the coordinator enters the blood pressure into a validated form at the point of care, and that entry is the first and only recording of the observation. There is no antecedent document. The record and the report are the same act. All three assumptions above fail at once.

This is why the change is larger than it looks. You have not simply removed a keystroke. You have removed the independent second record that a large part of clinical quality control was designed to exploit.

DimensionTranscription model (conventional EDC)Direct capture model
Where the source livesSite worksheet, paper chart, or electronic medical recordThe data acquisition tool itself, as declared in the protocol
Number of records per observationTwo, created at different timesOne, created once
Primary assurance activityCompare the eCRF against the source documentProve correct creation at the point of capture
Evidence a monitor examinesDocuments and the eCRF side by sideAudit trail, metadata, access log, delegation log
Query resolution basisConsult the source documentAudit trail plus a documented reason for change
Timing of edit checksAt data entry, days or weeks after the visitDuring the visit, while the participant is present
Who holds an independent copyThe site, by defaultNobody, unless the design deliberately creates one
Effect of a system outageData entry is delayed; source is unaffectedSource capture itself is blocked

Read the last two rows carefully. They are where most implementations get into trouble, and they are almost never covered in a vendor demonstration.

The Source Data Question Is the Whole Question

Every downstream consequence in this article follows from a single decision: which data elements are source, and where do they live. Under ICH E6(R3), that decision is not optional and not implicit. The investigator is required to define what is considered to be a source record, the methods of data capture, and their location before the trial starts, and to update that definition when needed.3 The same guideline adds that unnecessary transcription steps between the source record and the data acquisition tool should be avoided, which is the closest thing to regulatory encouragement for direct capture that currently exists.

E6(R3) goes further in its protocol content appendix. Among the data handling items a protocol must address is the identification of data to be recorded directly into the data acquisition tools, meaning data with no prior written or electronic record, and treated as the source record.3 That is a protocol-level declaration. It is reviewed, approved, and inspectable. If your team is treating the source data declaration as a data management plan detail to be resolved during startup, you have already put it in the wrong document.

Declaring the source before the first participant

The practical work here is a field-level inventory. For each variable the protocol collects, the team decides whether it originates in routine care documentation, in a trial-specific worksheet, in a device or laboratory system, or directly in the data acquisition tool. The FDA eSource guidance frames this as identifying the authorized originators of data, and it expects data element identifiers that allow an inspector to trace who or what created each element and when.1 The originator may be a person, a machine such as a wearable or sensor, or another computer system, a definition E6(R3) carries into its glossary.3

Doing this properly takes longer than teams expect, and it surfaces disagreements that were previously invisible. A vital sign captured by a monitor and read aloud to a coordinator has a different originator from one entered by a device interface. A concomitant medication history reconstructed from a hospital chart is not directly captured, no matter which screen it is typed into.

The mixed-source trial is the normal case, not the exception.

Very few studies are entirely direct capture. The realistic design has some fields declared as source in the data acquisition tool, some transcribed from the medical record, and some transferred from external systems. That mixture is acceptable, and E6(R3) explicitly contemplates verification effort scaled to data criticality when data captured on paper or in an electronic health record are manually transcribed into a computerized system. What is not acceptable is a monitoring plan and a data management plan that treat every field the same way. If your monitoring plan cannot state, field by field, whether verification applies, the plan is not implementable at the site.

What replaces verification

When the eCRF is the source, the assurance question changes from a comparison to a construction question. Nobody can tell you whether the recorded systolic pressure matches an independent record, because there is no independent record. What can be established is whether the record was created under conditions that make it trustworthy: by an authorized person acting under a documented delegation, at the time of the observation, in a validated system, with a complete audit trail, with the reason for any later change captured, and with the investigator retaining continuous access and control.

That is a different discipline. It is closer to computerized system assurance and data integrity practice than to conventional clinical monitoring, and it draws on skills that sit in quality and IT functions rather than in clinical operations. Organizations that treat direct capture as a clinical operations initiative and leave quality and IT out of the design tend to discover the gap during an inspection.

Audit Trails, Timestamps, and Attribution

If the audit trail is now the evidence that the source record is trustworthy, the audit trail has to be built to carry that weight. Most are not. The requirements below are specific, and they are the ones that separate a system suitable for direct capture from an EDC system that merely permits direct entry.

Field-level, not form-level

The EMA qualification opinion is unusually blunt on this point. Good clinical practice requires that all entries, changes, and deletions in a system are fully audit trailed, and in the case of eSource the audit trail should be per field. An audit trail at the end of a submitted form is not sufficient.2 Any change resulting from an automated data entry check must be visible as well, which means the system cannot silently normalize or correct an entry on the way to storage.

E6(R3) reinforces this from the sponsor side. Systems must be designed so that initial data entry and any subsequent changes or deletions are documented, including the reason for change where appropriate; audit trails, reports, and logs must not be disabled; and audit trails must not be modified except in rare circumstances such as inadvertent inclusion of personal information, and only with a logged justification.3

Timestamps that mean something

E6(R3) asks that the automatic capture of date and time of data entries or transfer be unambiguous, giving coordinated universal time as the example.3 This sounds like a technical footnote and is not. In a multi-country trial with tablets carried between clinics, local device clocks drift, time zones change mid-study, and daylight saving transitions create entries that appear to precede the events they describe. When the timestamp is the primary evidence of contemporaneous capture, an ambiguous timestamp is a data integrity finding waiting to be written.

Attribution and delegation

Attributable is the first letter of ALCOA+, and under direct capture it carries more weight than usual. The system must maintain a record of individual users authorized to access it, their roles, and their permissions, and E6(R3) requires that access permissions granted to investigator site staff be in accordance with delegations by the investigator and visible to the investigator.3 Shared accounts, generic site logins, and coordinator accounts left active after a staff member departs are common findings in electronic records generally. Under direct capture they do not simply weaken the record. They destroy the attribution of the source itself.

Investigator control of the only copy

This is the requirement most often missed, and the consequences of missing it are severe. E6(R3) states that the sponsor should not have exclusive control of data captured in data acquisition tools, in order to prevent undetectable changes, and that the investigator must have access to the required data for retention purposes.3 The EMA qualification opinion explains what is at stake. Missing continuous investigator control over eCRF data is already a frequent inspection finding. As long as sponsor-independent source data exist and an audit trail is possible, at least a verification of the eCRF data against the sponsor-independent source data can be carried out. The elimination of sponsor-independent source data, the opinion says, would significantly affect data integrity and therefore change the classification of these findings from major to critical.2

The severity escalation is the point.

A control weakness that would be graded major in a conventional trial can be graded critical in a direct capture trial, because there is no second record to fall back on. This is the single most important sentence in the EMA opinion for anyone building a business case. Direct capture raises the consequence of a control failure, which means the control design has to be better, not equivalent.

The practical answer is a certified copy strategy. Both workflow scenarios described in the EMA opinion end the same way: the direct data capture database transfers certified copies of the source data, and of the data reported to the sponsor if different, into the investigator trial master file before the site’s access to the sponsor system is removed.2 If your implementation plan does not name who produces those certified copies, in what format, at what trigger points, and how their completeness is verified, the plan is incomplete.

1

Inventory the audit trail content per system

Establish what the system actually records for each event type: creation, modification, deletion, workflow actions, user administration, and data transfer. Do this against the live configuration, not the vendor brochure. E6(R3) asks responsible parties to evaluate the system for the types and content of metadata available.

2

Decide which metadata are reviewed and retained

Reviewing everything is not achievable and is not asked for. E6(R3) expects the responsible party to determine which identified metadata require review and retention, with the extent and nature risk based and adjusted from experience during the trial.

3

Write the audit trail review procedure before enrollment

Name who reviews, at what frequency, against what triggers, and what happens when a pattern is found. An audit trail nobody reads is evidence of nothing. Under direct capture it is the primary quality control record for source data.

4

Prove investigator access, then prove it again at close out

Test read access, export capability, and the certified copy process during user acceptance testing, not at database lock. Verify that access survives the end of the trial and the end of the vendor contract.

5

Plan for the outage

Define the documented fallback when the tablet fails, the network drops, or the platform is unavailable during a visit. Paper capture followed by transcription is an acceptable answer if it is written down, controlled, and reflected in the source data declaration. An undocumented workaround invented at the bedside is not.

None of these requirements is exotic. Every one of them exists in some form under 21 CFR Part 11, which requires secure, computer-generated, time-stamped audit trails that independently record operator entries and actions, limits system access to authorized individuals, and requires that record changes not obscure previously recorded information.4 What direct capture changes is that these controls stop being administrative hygiene and become the only evidence that the clinical observation is real.

What Happens to Monitoring: Redistribution, Not Savings

The commercial case for direct capture usually leans on monitoring. The reasoning runs: source data verification is expensive, direct capture removes the need for it, therefore direct capture pays for itself. The first clause is true. The second is roughly true. The conclusion does not follow.

The evidence against generalized source data verification is genuinely strong

It is worth being clear that the case against comprehensive source data verification was settled well before direct capture became a practical option. TransCelerate BioPharma evaluated source data verification as a quality control measure through a literature review, retrospective analysis of trial data from member companies, and assessment of major and critical internal audit findings, and concluded that generalized source data verification has limited value as a quality control measure.5 An independent review published in the European Journal of Clinical Pharmacology analyzed 22 publications and found that fourteen showed little objective evidence of improved data integrity from traditional monitoring, including full source data verification, compared with reduced verification, central statistical monitoring, and remote monitoring. Eight publications identified potential for significant reductions in monitoring expenditure from reducing verification without compromising the validity of trial results. The authors concluded that full source data verification is not a rational method of ensuring data integrity and subject safety given the expense involved.6

The FDA reached a compatible position in its risk-based monitoring guidance, which supports monitoring approaches focused on critical data and processes rather than uniform verification of everything.7 So the direction of travel is not in dispute.

What does not disappear

Here is where the business case usually overreaches. E6(R3) defines monitoring as a broad range of activities including communication with investigator sites, verification of investigator and site staff qualifications and site resources, training, and review of trial documents and information using a range of approaches including source data review, source data verification, data analytics, and visits to facilities undertaking trial-related activities.3 Source data verification is one item in that list. Removing it leaves the rest intact, and in a direct capture trial several of the remaining items get larger.

GROWS

Source data review

Verification asks whether the eCRF matches the document. Review asks whether the record itself is complete, internally consistent, and medically coherent. That question survives direct capture untouched, and it requires clinical judgment rather than comparison, so it is harder to delegate and harder to automate.

GROWS

Centralized statistical monitoring

E6(R3) treats centralized monitoring as an evaluation of accumulated data and notes that these activities may be conducted by persons in different roles, giving data scientist as the example. That is a hiring statement, not a workflow note.

GROWS

Audit trail and metadata review

New activity with no equivalent in the transcription model at anything like the same scale. Somebody has to design it, staff it, and document it, and the reviewer needs enough system knowledge to interpret what they are seeing.

GROWS

Site qualification and system oversight

The EMA opinion expects a site qualification procedure before deploying the system at any given site, plus specified help desk availability and accessibility. Sponsors also carry the validation burden, including study-specific validation and validation of transfers into the site’s medical record.

Central statistical monitoring is worth dwelling on, because it is where the displaced effort most usefully lands. The methods are mature. Key risk indicator approaches have been developed and applied within large ongoing randomized trials, monitoring factors such as serious adverse event reporting rates, treatment compliance, and laboratory results to focus attention on variables most likely to affect reliability or participant safety.8 Unsupervised statistical monitoring using mixed effects models has been shown to detect centers with fraud or other anomalies. In a reanalysis of the ESPS2 trial, five centers were flagged as atypical, the center with known fraud ranked second among them, and an incremental analysis showed that the fraudulent center could have been detected after only 25 percent of its data had been reported.9 Central statistical monitoring has also been applied successfully to investigator-led oncology trials, where budgets rarely permit extensive on-site verification.10

The important observation for a direct capture program is that these techniques find a class of problem that source data verification never found. Fabrication, unusually low variability in vital signs and laboratory values, absent adverse events, and implausible distributions are visible in aggregate and invisible in a single participant file. The shift is therefore not a downgrade of quality control. It is a change of instrument.

22 Publications reviewed in the European Journal of Clinical Pharmacology analysis; 14 showed little objective evidence that traditional monitoring improved data integrity over reduced verification6
25% Proportion of a fraudulent center’s data that had been reported by the time unsupervised statistical monitoring could have detected it in a reanalysis of the ESPS2 trial9
5.9M Average datapoints collected per Phase III protocol, growing at 10.8 percent annually, in the Tufts CSDD and TransCelerate analysis of 105 protocols11

Present this to a steering committee as a redistribution and you will keep the program’s credibility. Present it as a saving and you will spend the second year explaining why the monitoring line did not fall as promised, while the data management and statistics lines quietly grew.

The Site Burden Reality

Direct capture is almost always sold as reducing work at the site. Sometimes it does. Frequently it does not, and the reason is structural rather than a matter of poor execution.

The industry position on this is not naive. TransCelerate’s published point of view on accelerating eSource adoption, which groups the field into four types (electronic health records, devices and apps, non-CRF data transfers, and direct data capture), states plainly that eSource should not increase participant or site burden and identifies duplicate data entry and inadequate training as ways in which implementations do exactly that. The same analysis lists interoperability gaps between healthcare and research systems, unclear processes for correcting source data, and gaps in eSource and informatics expertise among the barriers holding adoption back.18 None of those barriers is solved by choosing a better capture screen.

The dual entry trap

A research site is usually also a care setting. The institution has a legal and professional obligation to maintain a medical record for the participant’s standard of care, and that obligation is independent of the trial. If protocol-mandated observations are captured in the sponsor’s direct capture system and that system does not write back into the institution’s medical record, the site staff enter the same information twice. The transcription step you removed from the sponsor’s side has reappeared on the institution’s side, where nobody is measuring it.

The EMA qualification opinion recognizes this and is direct about the obligation it creates. To decrease workload on the investigator and site staff and to avoid transcription errors, transcription requiring manual intervention between eSource and the medical record should be avoided. Using an eSource must not result in a depletion or disorder of the information available in patients’ medical records. An increase of the investigator staff’s workload must be avoided. The opinion states that the long-term ambition should be that collected data can be transferred automatically into the site’s own electronic medical record, or captured automatically from it, and that cooperation to achieve standardization of data interoperability should be supported.2

The opinion also names the failure mode that scales worst. Investigators must use different eSource systems for the various trials conducted by different sponsors and vendors in parallel. If those systems are not compatible for data transfer into the medical records, the result is increased data dispersion, depleted medical records, increased workload for site personnel, and a potential breach of national requirements for the upkeep of medical records.2 A single sponsor optimizing its own trial can make the site worse off overall. Nothing in a single-sponsor business case will surface that.

Questions to ask before a direct capture program is approved.
  • Does the system write captured data back into the site’s medical record, and through what validated mechanism? If the answer is a certified copy uploaded as a document, is that acceptable to the institution and readily traceable within its record?
  • What does the site do about protocol data that must appear in the medical record for continuity of care, when the sponsor is only permitted to receive pseudonymized data?
  • How many other sponsors’ systems is this site already running, and who at the site has counted the total training and login burden?
  • Who pays for the site time spent on entry into a second system, and does the site budget recognize it?
  • What happens to site access to the data after the trial closes and the vendor contract ends?

What actually reduces burden

The honest answer is that the largest lever is not capture technology. It is collecting less. The Tufts Center for the Study of Drug Development and TransCelerate analyzed 105 protocols and found that Phase III protocols now collect an average of 5.9 million datapoints, a 67 percent increase since 2020 and 10.8 percent annual growth. Roughly one third of procedures and data collected per Phase III protocol were classified as non-core or superfluous, and between 25 and 30 percent of participant and site burden was associated with non-core or non-essential procedures.11 Questionnaires carried the highest concentration of non-core and non-essential content.

Set that alongside a direct capture program and the arithmetic is uncomfortable. Removing a transcription step from a protocol that collects a third more data than it needs is optimizing the wrong layer. Protocol simplification and direct capture are complementary, and the order matters: simplify first, then decide how to capture what remains.

Where site burden genuinely falls.

Two situations produce real reductions rather than displaced work. The first is trial-specific data that has no home in routine care documentation, where investigators today create their own paper worksheets and then transcribe them. The EMA opinion identifies exactly this case and suggests that replacing such worksheets with sponsor-provided electronic worksheets is likely to improve data quality, giving rating scales not used in normal clinical practice and detailed recording of multiple blood sampling times as examples. The second is validated automated transfer from the site’s own systems, where the source remains in the institution’s record and the transfer, rather than the transcription, becomes the controlled step. When transfer from the medical record is automated, source data verification forms part of the system validation rather than a site visit activity.2

Both of those work because they remove a genuine duplicate act. What does not work is moving the point of entry from one screen to another screen and calling it a reduction. Site coordinators can tell the difference immediately, and their assessment tends to be more accurate than the projected savings in the business case.

Queries and Data Cleaning Without a Document to Consult

Query management is where the absence of a source document becomes concrete for the data management team, and it is the area most often left until the study is already enrolling.

The queries that never happen

Some of the change is favorable and immediate. In the transcription model, edit checks fire when the data reaches the eCRF, which may be days or weeks after the visit, by which time the participant has gone home and the coordinator is reconstructing what happened. In direct capture, validation runs while the data is being entered during the visit. The EMA opinion notes that edit checks configured by the sponsor take place when the data is entered in the system and may help reduce or identify missing or erroneous entries, and E6(R3) asks that automated data validation checks at the point of capture be considered based on risk, with their implementation controlled and documented.23

The practical effect is that a large share of range checks, missing field prompts, and simple consistency failures are resolved at the moment of capture and never become queries at all. That is a genuine improvement in cycle time and in data quality, because the correction is made by the person who made the observation while the observation is fresh.

The queries that remain, and what they resolve against

What remains is the harder category: cross-form inconsistencies, clinical plausibility concerns, protocol interpretation questions, and coding issues. In the transcription model, these are resolved by consulting the source document. Under direct capture there is no such document, so the resolution has to rest on something else.

E6(R3) is specific about what that something else looks like. Corrections should be attributed to the person or computerized system making the correction, justified, supported by source records around the time of original entry, and performed in a timely manner.3 When the entry is the source record, the support for a correction is the audit trail plus a documented, meaningful reason for change. This has three consequences that data management leads should plan for.

First, reason for change stops being an optional field. In many conventional studies, reason-for-change text is either not required or is filled with a default value. Under direct capture it is the substantive record of why the source was altered, and generic entries such as “data entry error” carry no evidentiary weight when an inspector asks how a value changed three weeks after a visit.

Second, the sponsor’s latitude to correct data narrows sharply. E6(R3) states that the sponsor should not make changes to data entered by the investigator or trial participants unless justified, agreed upon in advance by the investigator, and documented, and that the sponsor should allow correction of errors where requested by investigators or participants.3 Practices such as self-evident corrections, where data management fixes obvious errors under a standing agreement, need rewriting for a study where those data are the source record.

Third, the query routing itself becomes an engineering problem. Both workflow scenarios in the EMA opinion describe queries raised in the sponsor’s clinical database and passed back through the direct capture database to the capture tool, where the site resolves them.2 That round trip crosses at least two validated systems and a mapping layer. Every hop needs to preserve the audit trail, and the mapping from the capture database into the eCRF database has to be performed via a validated process.

Data cleaning activityTranscription modelDirect capture model
Range and format errorsQuery raised after entry, resolved by site days laterPrevented at the point of capture during the visit
Value disputed by monitorCompare eCRF against worksheet or chartExamine audit trail, entry time, and user; escalate to investigator judgment
Correction after entryCorrect eCRF to match source; source unchangedCorrect the source itself; audit trail and reason for change are the only record of the prior state
Reason for changeOften administrativeSubstantive evidence, reviewed and inspectable
Sponsor-initiated correctionCommon under standing agreementsRequires prior investigator agreement and documentation
Close out evidenceSite retains its own source documentsCertified copies transferred to the investigator file before access removal

None of this is difficult in principle. It is difficult in practice because the artifacts that govern it, the data management plan, the monitoring plan, the query conventions document, and the site training materials, were all written for a world with two records and have to be rewritten for a world with one.

Where Direct Capture Fits and Where EDC Still Makes Sense

The useful question is not whether direct capture is better. It is which data elements, in which studies, at which sites, benefit from it. Three variables drive the answer.

Study type and where the data originates

Direct capture works best where the observation does not already exist in the medical record and would otherwise be written on a trial worksheet. The EMA opinion gives the clearest examples: rating scales not used in normal clinical practice, detailed recording of multiple blood sampling times, and similar protocol-driven parameters. It also suggests that trials already using electronic technology such as electronic patient-reported outcomes, eCRFs, real-time outcome monitoring, and electronic capture of laboratory results are a reasonable initial testing ground.2

Conventional EDC remains the sensible model where the observation originates in routine care and is properly documented in the institution’s record. Oncology response assessment, complex safety histories, prior treatment records, and anything a treating physician records for care purposes belong in the medical record first. Trying to make the trial system the source for those elements risks exactly the depletion of the medical record that the EMA opinion warns against.

Endpoint complexity and the amount of free text

Structured, well-defined endpoints suit direct capture. Endpoints requiring narrative clinical reasoning do not, or at least not without careful design. The EMA opinion notes that the capture tool should not be limited to structured data only and must allow free text, and that use should be evaluated by in-use testing of the eSource approach against collecting the same data without it, to confirm that the technology does not degrade the interaction between investigator and participant.2 That in-use comparison is rarely performed and is one of the cheapest risk reduction steps available.

A further constraint follows from the certified copy requirement. The electronic worksheet should only contain elements that can be adequately mirrored in a printout or flat file without loss of information, because a certified copy must be producible.2 Highly interactive form designs with conditional logic, embedded calculations, and dynamic content can fail this test.

Site infrastructure and institutional policy

The third variable is the one most often assumed away. Direct capture depends on connectivity, device availability, and institutional willingness to accept that some care-relevant information is first recorded in a sponsor-provided system. Academic medical centers with strict record-keeping policies, sites in countries with prescriptive medical record legislation, and sites already running several sponsor systems are poor candidates regardless of how attractive the protocol looks. The EMA opinion expects a site qualification procedure before deploying the system at any given site, and that qualification should genuinely be able to return a negative answer.2

GOOD FIT

Protocol-only assessments

Rating scales, structured symptom instruments, sampling time logs, and other data that exist solely because the protocol requires them and would otherwise live on a paper worksheet.

GOOD FIT

Participant and device generated data

Electronic clinical outcome assessments, wearables, and sensors, where the data acquisition tool has always been the source and the model is already understood by regulators and sites.

POOR FIT

Care-derived endpoints

Observations a clinician records for treatment purposes and which must remain complete and accessible in the institutional record. Here the better investment is validated transfer out of the medical record, not capture into a separate system.

POOR FIT

Sites without the operating conditions

Limited connectivity, no dedicated research staff, institutional policy that forbids first recording outside the medical record, or a site already carrying several sponsor systems.

A realistic target for most sponsors today is a hybrid: direct capture for a defined subset of protocol-specific fields, validated transfer for laboratory and device data, and conventional transcription with risk-proportionate verification for care-derived elements. That is not a compromise position. It is what the source data declaration will produce if the field-level inventory is done honestly.

Regulatory Expectations You Should Plan Against

The regulatory position on direct capture is more settled than most internal debates suggest. The gap is not in the guidance. It is in how few implementation plans map to it.

ICH E6(R3) is the anchor

ICH adopted E6(R3) in January 2025, and the FDA issued its final guidance adopting the revised guideline in September 2025.312 Three features matter most for a direct capture program. The guideline requires the investigator to define source records, capture methods, and locations before the trial begins. It requires the protocol to identify data recorded directly into data acquisition tools with no prior record and treated as source. And it builds an entire section on data governance across the life cycle covering capture, metadata and audit trails, review, corrections, transfer, finalization, retention, and destruction.3 A program that can demonstrate compliance with that section is, in practice, a program that has answered most of the questions in this article.

FDA expectations

The 2013 eSource guidance remains the most specific United States statement on the subject, addressing identification of authorized data originators, data element identifiers that support examination of the audit trail, capture of source data into the eCRF, and investigator responsibilities including review before signing.113 The 2018 guidance on the use of electronic health record data in clinical investigations covers the adjacent case where the institution’s record feeds the trial, including expectations for interoperability and for the sponsor’s assessment of the source system.14 The 2024 final guidance on conducting clinical trials with decentralized elements addresses the operating context in which direct capture is most often deployed, including validation and reliability of the technologies used for data collection.15 Underneath all of them, 21 CFR Part 11 governs the electronic records and signatures themselves.4

EMA expectations

The 2019 qualification opinion on eSource direct data capture remains the most detailed European treatment of the specific question and is worth reading in full by anyone designing a program.2 It was issued in the knowledge that a broader guideline on computerized systems and electronic data in clinical trials was under development, and the agency noted at the time that the guideline would constitute the definitive guidance once in force. The opinion is not general guidance, and it was written against one applicant’s proposal, but its reasoning on investigator control, certified copies, per-field audit trails, medical record depletion, and finding severity is the clearest published statement of what European inspectors will look for.

What inspectors already find

It helps to know where problems concentrate today. MHRA’s published good clinical practice inspection metrics show that case report form and source data issues have been among the most common major findings at investigator site inspections, alongside delegation of responsibilities, protocol compliance, and data integrity.16 Those are precisely the areas that direct capture makes more consequential rather than less. Delegation errors become attribution errors. Source data location ambiguity becomes an unanswerable question. Data integrity weaknesses lose the fallback of a second record.

Pilot experience supports a measured pace. A hospital-based pilot in China constructing and evaluating an electronic source data flow from hospital electronic medical records reported that the approach was feasible while identifying substantive quality and standardization work required to make the source data usable for trial purposes.17 The lesson from published pilots is consistent: the technology works, and the surrounding process, standards, and governance work is where the effort goes.

A workable sequence for a first program

Choose one study with a well-defined set of protocol-only assessments. Perform the field-level source inventory and write the protocol declaration. Design the certified copy and investigator access mechanism before selecting a vendor, so it becomes a selection criterion rather than a gap discovered later. Qualify a small number of sites with the operating conditions the design assumes, and be prepared to exclude sites that do not have them. Rewrite the monitoring plan, data management plan, and query conventions for a one-record world, and train sites on the corrections and reason-for-change discipline specifically. Then measure the site time actually spent, not the site time projected, and use that number for the next study.

Conclusion

Direct data capture is a small technical change wrapped around a significant change in the evidentiary structure of a clinical trial. Removing transcription removes a source of error, which is a real benefit. It also removes the independent second record that source data verification, query resolution, and a good deal of institutional habit depend on. The organizations that struggle with direct capture are rarely the ones that picked the wrong platform. They are the ones that treated the source data declaration as paperwork, assumed monitoring effort would simply fall away, and did not ask what the site would do about its own medical record. The organizations that succeed treat it as what it is: a data governance change that happens to involve a new screen at the point of care.

Our advice to clients considering this is consistently the same. Do the field-level source inventory first, before any vendor conversation, because it tells you how much of your protocol is actually a candidate. Model the monitoring change as a redistribution across roles and functions, not a reduction, and get the receiving functions staffed before enrollment. Ask the sites what happens to their medical record, and believe their answer. And design the certified copy and investigator access mechanism at the start, because it is the control that determines whether a future finding is graded major or critical.

Sakara Digital works with pharma and biotech organizations designing clinical data models that hold up under inspection. If you are weighing direct data capture against conventional EDC for an upcoming program and want an independent read on where it fits and what it will actually require, we are happy to have that conversation.

For Further Reading