Anyone comparing AI-assisted and manual ICSR processing is usually building a business case, and business cases in this area fail for a predictable reason: they assume automation applies evenly across the workflow. It does not. Some steps compress dramatically, some compress a little, and some do not compress at all — and the ones that do not are often the ones consuming your most expensive people.
This guide goes step by step through the case lifecycle, says honestly what changes at each, compares how the two approaches fail, and gives you a method for measuring the difference on your own case mix rather than adopting anyone's benchmark.
No invented benchmarks here — and a disclosure
We build a platform with AI-assisted features, so treat our conclusions accordingly. You will not find percentage improvement figures in this guide, because any such number depends entirely on case mix, source document quality, existing process maturity and how the baseline was measured — which means a vendor's figure tells you about their pilot, not your operation. What follows instead is a method for producing the number for yourself.
Before comparing anything, it is worth being precise about what manual processing consists of. Teams often describe it as "data entry and review", which hides where effort concentrates.
| Step | Nature of the work | Who typically does it |
|---|---|---|
| Intake and logging | Administrative: retrieving from a mailbox, saving source documents, creating a record | Case processor / administrator |
| Duplicate search | Repetitive searching across partial identifiers, several attempts per case | Case processor |
| Reading the source | Comprehension — unavoidable and genuinely valuable | Case processor |
| Transcription into fields | Pure keystrokes: patient, product, dose, dates, events, reporter, labs | Case processor |
| Coding | Dictionary lookup, synonym resolution, term selection | Case processor / coder |
| Narrative writing | Composition from the structured data just entered | Case processor |
| Quality review | Field-by-field comparison against source, plus consistency checks | QC reviewer |
| Medical assessment | Clinical judgement: seriousness, causality, expectedness | Physician / qualified HCP |
| Reportability and deadlines | Determining destinations and dates, often against a reference document or spreadsheet | Case processor / regulatory |
| Submission and ACK handling | Generating output, transmitting, checking acknowledgements | Case processor / regulatory |
The uncomfortable observation
In most operations the majority of case-processing hours go to transcription, duplicate searching, coding lookup, narrative composition and field-by-field QC — none of which requires pharmacovigilance training. The steps that genuinely need qualified judgement are a minority of the elapsed time and are frequently the ones done last, under deadline pressure, by the most expensive person in the chain.
Honest assessment of each step. Note how many rows say "unchanged" — that is the point of the exercise.
| Step | Manual | AI-assisted | Realistic change |
|---|---|---|---|
| Intake and logging | Manual retrieval and record creation | Channel-routed, auto-created, source attached | Large reduction — mostly administrative work removed |
| Duplicate search | Manual searches on partial identifiers | Scored candidates surfaced automatically | Large reduction in effort; the confirm/reject decision remains human |
| Reading the source | Reviewer reads the document | Reviewer still reads it, to verify | Unchanged — and should be |
| Transcription | Every field typed | Fields extracted with per-field confidence; reviewer verifies | Largest single reduction available |
| Coding | Manual dictionary lookup and selection | Terms proposed; coder accepts or overrides | Substantial reduction; clinical appropriateness still reviewed |
| Narrative | Written from scratch | Draft generated; reviewer edits | Moderate reduction — editing is faster than composing |
| Quality review | Full field-by-field comparison | Mechanical checks automated; effort focused on low-confidence fields and narrative-versus-data agreement | Moderate reduction; the comprehension check is unchanged |
| Medical assessment | Clinical judgement | Clinical judgement | Unchanged in substance — but the assessor receives a cleaner, faster-arriving case |
| Reportability and deadlines | Manual determination and tracking | Rules engine derives destinations, Day 0 and deadlines | Large reduction, and a step change in reliability |
| Submission and ACK | Manual generation, transmission, ACK checking | Generated, validated, transmitted, ACK reconciled to the case version | Large reduction, with the silent-failure risk closed |
What does not change
Reading the source document, the narrative-versus-data comprehension check, and every regulated determination — seriousness, causality, expectedness, release. A business case that assumes these compress will overstate the benefit and, worse, may lead to a process design that quietly weakens the controls those steps represent.
This is the part most comparisons skip, and it matters more than raw speed for anyone thinking about quality risk. Manual and AI-assisted processing fail in different ways, which changes where you should place your controls.
| Manual processing | AI-assisted processing | |
|---|---|---|
| Typical errors | Transposition, omitted fields, inconsistent coding between processors, fatigue effects late in a shift | Misread values from poor-quality sources, plausible-looking but wrong extraction, systematic misinterpretation of an unusual document layout |
| Error distribution | Random and individual — varies by person, day and workload | Systematic and correlated — if it misreads one layout it misreads all of them |
| Detectability | Hard to predict which cases are affected | Largely predictable via confidence scoring, which is the mitigating advantage |
| Volume sensitivity | Error rate tends to rise with workload and time pressure | Error rate is stable regardless of volume |
| Control that works | Sampling, second review, training, workload management | Confidence-based routing, source-context verification, monitoring for systematic drift |
The practical implication
Systematic errors are more alarming in principle but easier to catch in practice, because they are consistent and confidence scoring flags the uncertain cases for you. Random human errors are individually smaller and harder to target. The right response is not to prefer one but to place different controls: confidence-based routing plus drift monitoring for the automated path, sampling and workload management for the manual one.
Comparisons often stall on licence price, which is the wrong denominator. The number that decides the business case is fully loaded cost per case — and it has to include the parts that do not appear on an invoice.
- Processing labour by role, since a physician hour and a processor hour are not interchangeable costs
- Rework — cases returned from QC or medical review, and the cost of a correction after submission
- Avoidable rejections from an authority, which consume clock as well as labour
- Late-submission consequences, which are compliance costs rather than operational ones
- Platform, hosting and gateway licensing, noting whether an AS2 gateway is included or separate
- Dictionary licences (MedDRA, WHO-DD), which are the client's responsibility in essentially all arrangements
- Qualification and re-qualification effort, including after platform changes
- The cost of capacity you cannot flex — the overtime or contractor spend when volume spikes
One line item that is always yours
Third-party dictionaries and licensed content - including MedDRA and the WHO Drug Dictionary - are procured and licensed by the client. PVgenix integrates them into the application.
For a fuller breakdown of the cost components across system categories, see how much does pharmacovigilance software cost.
Rather than adopting a vendor figure, produce your own. This is a measurable exercise and it is the only version of the number that will survive scrutiny internally.
- **Establish a baseline properly.** Measure current per-step effort across a representative sample — at least one full month, covering your real mix of serious and non-serious, initial and follow-up, clean and messy source documents. Record elapsed time and touch time separately.
- **Segment by source document type.** Structured web form, typed email, clean PDF, scanned document, handwritten fax. Improvement varies enormously across these, and a blended average hides which of your channels actually benefits.
- **Define the quality bar first.** Agree what accuracy and completeness you require before piloting, so you are not tempted to trade quality for speed retrospectively.
- **Pilot on real historical cases.** Run cases you have already processed through the AI-assisted path and compare field by field against the known-good record. This gives you a true accuracy measure rather than a demo impression.
- **Measure the review effort, not just the extraction.** The relevant figure is total time to a released case, including verifying low-confidence fields — not time to a populated form.
- **Track error types, not just an error rate.** You need to know whether misses are random or systematic, because that determines your controls.
- **Include the deadline metrics.** On-time submission rate and backlog ageing often improve more than raw cycle time, and they matter more.
- **Re-measure after the learning period.** Early numbers reflect unfamiliarity with the review console as much as anything else.
Do not pilot only on clean cases
The most common way a pilot misleads is by using tidy, well-structured source documents. Your real inbound mix includes the scanned fax and the three-line email with an ambiguous date. Deliberately include the awkward material, because that is where the review effort concentrates in production.
- Very low case volume — if you process a small number of cases a month, the constraint is compliance coverage and audit trail quality, not throughput, and the business case should be built on those
- Serious and fatal cases — these should be processed at full depth regardless of what automation is available
- Any determination carrying regulatory accountability, as covered in AI in pharmacovigilance
- A broken process — automating an undefined workflow produces faster inconsistency; fix the process first
- As a substitute for capacity you actually need — automation redirects review effort, it does not remove the requirement for qualified people
In practice the sensible end state is not "AI processing" versus "manual processing" but a single workflow that routes by risk and confidence — which is also the configuration you can defend at an inspection because it was decided rather than inherited.
| Case profile | Path |
|---|---|
| Serious, fatal, or expedited-reportable | AI-assisted extraction and coding, then full human QC and medical review |
| Any case with low-confidence extracted fields | Mandatory field verification against source regardless of seriousness |
| Non-serious, listed, established product, high-confidence extraction | Candidate for configured straight-through processing within documented scope |
| Clinical trial cases, potential SUSARs | Full depth — expedited consequences and blinding considerations warrant it |
| Poor-quality source documents | Routed for human-led entry rather than forced through extraction |
| Follow-up with potentially significant information | Human significance determination, always |
In PVgenix, extraction is confidence-scored per field with automatic routing of low-confidence items to review, straight-through processing is configurable and scoped to qualifying low-risk case types with thresholds you set, and reportability and deadlines are derived by a rules engine rather than inferred. See AI case intake for the extraction and review console and AI pharmacovigilance software for how the boundary is enforced.
The short version
The gains are real and concentrated in transcription, duplicate searching, coding lookup, narrative drafting, reportability derivation and submission handling. They are approximately zero in reading the source and in every regulated determination. Build the business case on cost per case rather than licence price, measure it on your own historical cases including the messy ones, and expect review effort to be redirected rather than removed.
Also in this cluster: AI in pharmacovigilance for the boundary between assistance and decision, and AI MedDRA coding explained for the mechanics of the most-automated single step.
Frequently asked questions
Common questions
It depends so heavily on case mix, source document quality and existing process maturity that any single figure is misleading — which is why we do not publish one. The gains concentrate in transcription, duplicate searching, coding lookup, narrative drafting, reportability derivation and submission handling. They are close to zero in reading the source document and in the regulated determinations. Measure it on your own historical cases, segmented by source document type, and include review effort rather than measuring only extraction.
Reading and comprehending the source document, the narrative-versus-data consistency check, and every regulated determination: seriousness confirmation, causality assessment, expectedness or listedness, case validity, duplicate merge decisions, and release to submission. A business case that assumes these compress will overstate the benefit and risks a process design that weakens the controls those steps represent.
The more useful framing is that the error profiles differ. Manual errors tend to be random and individual — transposition, omissions, coding inconsistency between processors, fatigue late in a shift — and the rate tends to rise with workload. AI errors tend to be systematic and correlated: a layout misread once is misread consistently. Systematic errors sound worse but are easier to catch, because confidence scoring flags uncertain cases and drift monitoring detects patterns. The right response is different controls, not a preference.
Run real historical cases you have already processed through the AI-assisted path and compare field by field against the known-good record — that gives a true accuracy measure rather than a demo impression. Establish a proper baseline first over at least a month of representative volume, segment results by source document type, define your quality bar before you start, measure total time to a released case including verification, track error types rather than just a rate, and deliberately include awkward source documents. A pilot run only on clean cases will overstate the benefit.
Processing labour broken down by role, since physician and processor hours are not interchangeable costs; rework from QC and medical review returns; avoidable authority rejections, which consume clock as well as labour; late-submission consequences; platform, hosting and gateway licensing, noting whether AS2 is included or separate; MedDRA and WHO Drug Dictionary licences, which are always the client's responsibility; qualification and re-qualification effort; and the cost of inflexible capacity such as overtime or contractors when volume spikes. Licence price alone is the wrong denominator.
At very low case volumes, where the real constraint is compliance coverage and audit trail quality rather than throughput. For serious and fatal cases, which warrant full-depth processing regardless. For any determination carrying regulatory accountability. Where the underlying process is undefined — automating that produces faster inconsistency. And as a substitute for capacity you genuinely need, since automation redirects review effort rather than removing the requirement for qualified people.
