1. Translating Clinical Intent into Executable Logic
The initial clinical definitions were expressed through expert knowledge, published evidence, narrative criteria, care-setting requirements, laboratory thresholds, treatment patterns, and temporal relationships. These descriptions were not directly executable. Each phenotype required an explicit definition of the index event, eligibility criteria, observation requirements, temporal windows, exclusions, cohort exit, and recurrent-event handling. A clinically plausible description could still produce an analytically incorrect cohort if these elements were implemented inconsistently.
2. Combining Evidence Across Multiple OMOP Domains
The target clinical states could not be represented reliably by diagnosis codes alone. Relevant evidence was distributed across conditions, treatments, encounters, laboratory results, and other structured clinical events. The cohort logic therefore had to combine multiple OMOP domains while preserving the correct temporal and contextual relationships between them.
3. Building Complete and Clinically Precise Concept Sets
Terminology selection was more complex than searching for matching clinical terms. Errors clustered into three root causes, shown above: scope, concept identity, and mapping or domain assignment.
Scope and concept-identity errors were usually invisible until a clinical reviewer read through sample records, since the vocabulary looked correct on its face. Mapping and domain errors, by contrast, often surfaced only once a concept set ran against real data across institutions — a source-to-standard gap or an incorrect domain assumption that produced no visible error until execution exposed it. The client needed concept sets that were clinically defensible, technically compatible with OMOP, and sufficiently documented for review and reuse.

4. Supporting Execution Across Heterogeneous Data Sources
OMOP provided a common analytical structure, but data availability and representation still varied among institutions. These differences clustered into three root causes: coding and mapping practices, data completeness, and site-specific ETL process.
Data-completeness gaps often produced silent failures — a phenotype simply returned fewer results, with no visible error. Coding, mapping, and ETL differences typically surfaced as visibly incorrect results. Successful execution at one institution therefore did not guarantee equivalent results across the network.

5. Establishing Reproducibility and Governance
The client needed more than a set of queries. Each phenotype had to become a governed analytical asset with a documented clinical intent, accountable ownership, versioned concept sets, executable logic, quality-assurance evidence, review history, and change control. Without these controls, cohort definitions could not be reliably maintained, compared, or reused.
6. Integrating LLM Assistance Without Losing Expert Control
Large language models could accelerate evidence synthesis, terminology exploration, documentation, and logical review. However, unconstrained model output could introduce invented terminology, incomplete criteria, incorrect temporal assumptions, or unjustified clinical conclusions. The workflow therefore had to use LLMs selectively while keeping all clinically consequential decisions under human control.
1) Structured Phenotype Specification Framework
SciForce introduced a standardized specification template for converting clinical questions into implementable phenotype requirements. Each specification documented eleven components, grouped into the event defining cohort entry and the criteria determining analytical eligibility around that event, shown below.

Narrative clinical criteria could sound plausible and still produce an analytically incorrect cohort if these two groups were implemented inconsistently. The template's separation gave reviewers a fixed checkpoint: whether a patient qualified for cohort entry, and separately, whether they remained eligible under the criteria evaluated afterward.
2) OMOP Concept-Set Engineering
We created curated concept sets supporting selected critical-care phenotype specifications. Concept sets were developed as reusable semantic assets rather than as undocumented components embedded within individual queries. The curation process included:
Final concept membership remained subject to expert terminology and clinical review.
3) Executable Cohort Development
The approved phenotype specifications and concept sets were translated into executable OMOP cohort definitions developed in OHDSI ATLAS. ATLAS carried forward the entry and eligibility structure shown above, and added four implementation-specific mechanics: event-level restrictions, occurrence counts, incident and prevalent interpretations, and era construction. ATLAS definitions supported both human-readable review and export for execution.
One phenotype was implemented as a derived analytical construct built on top of its parent cohort, rather than as its own standalone definition, keeping primary cohort identification separate from downstream derivation.
4) Human-Supervised LLM Assistance
The workflow incorporated an LLM as a controlled knowledge-engineering and quality-assurance layer. The LLM supported:
The model did not independently approve concepts, determine clinical validity, or deploy production cohort logic. Final decisions remained with clinical, vocabulary, methodological, and technical reviewers.
5) Multi-Layer Quality Assurance
SciForce implemented a multi-layer quality assurance framework covering terminology, cohort logic, population-level behavior, and cross-source consistency. Each phenotype was evaluated using OHDSI CohortDiagnostics to assess cohort size, inclusion-rule attrition, index-event concept utilization, demographic and temporal distributions, observation time, visit context, and other indicators of data fitness and phenotype behavior.
The review also examined concept-set scope and hierarchy expansion, standard and non-standard concept use, source-to-standard mappings, domain alignment, index-event implementation, entry-versus-inclusion semantics, temporal and observation requirements, episode construction, cohort exit and censoring, and the consistency of measurement units and value representations.
Successful cohort generation was treated as a necessary technical check, not as evidence that the phenotype was clinically valid.
6) Cross-Source Diagnostic Framework
The cohort definitions were evaluated using OHDSI CohortDiagnostics across participating institutional OMOP data sources where the required data were available. Diagnostic outputs were reviewed both within individual sites and comparatively across sites to assess differences in cohort generation, inclusion-rule attrition, concept utilization, data availability, and temporal behavior. The cross-source review was designed to identify:
Cross-source variation triggered investigation into phenotype behavior and data fitness before any implementation error was assumed.
7) Governance and Documentation
Each phenotype was supported by documentation covering:
The resulting governance model supported revision, release, and reuse as the phenotype library grew.
Reusable OMOP Concept Library
The project delivered curated concept sets that could be reviewed, versioned, and reused across multiple cohort definitions. This reduced dependence on one-off code lists and established a shared semantic foundation for future phenotype development.
Executable and Human-Readable Cohort Definitions
The solution connected clinical descriptions with executable OMOP cohort logic. Researchers could review both the intended clinical meaning and its technical implementation instead of relying on opaque SQL or undocumented ATLAS configurations.
Multi-Domain Clinical Logic
The framework supported phenotype definitions based on combinations of conditions, treatments, visits, measurements, and temporal relationships. This enabled the representation of complex clinical states that could not be identified through a single code or domain.
Human-Governed LLM Workflow
The LLM was used as a human-supervised knowledge-engineering and quality-assurance assistant rather than as an autonomous phenotype generator. It supported evidence synthesis, converted narrative clinical definitions into structured phenotype specifications, expanded terminology searches, checked consistency between approved phenotype intent and ATLAS logic, drafted concept-set rationales and change documentation, and summarized cross-site diagnostic findings. Final concept selection, temporal logic, cohort implementation, and approval remained under expert control.
Cross-Database Diagnostic Readiness
The phenotype package was structured for systematic evaluation across heterogeneous OMOP data sources. Site-level differences in cohort counts, attrition, concept use, and data availability could be detected and investigated using a consistent diagnostic framework.
Versioned and Auditable Assets
Concept sets, phenotype specifications, cohort definitions, QA findings, and changes were documented as maintainable project assets. This made it possible to distinguish semantic changes from technical corrections and documentation-only updates.
Support for Primary and Derived Phenotypes
The architecture distinguished executable cohorts from analytical derivations and subphenotypes. This avoided forcing every clinical construct into the same implementation pattern and enabled more appropriate downstream analytical logic.
Privacy-Conscious LLM Integration
The LLM layer was designed to operate on non-identifiable artifacts such as literature, terminology metadata, phenotype specifications, cohort-definition files, and aggregate diagnostics. Patient-level cohort generation and analysis remained within secured institutional data environments.
1. Clinical Requirements Analysis
At first, we identified the intended use of each phenotype and resolving ambiguities in the clinical description. Clinical criteria were decomposed into:
This prevented premature concept selection before the clinical construct had been adequately defined.
2. Evidence and Data-Requirement Review
Relevant publications, clinical definitions, terminology resources, and available OMOP data elements were reviewed. We identified which clinical signals could be represented reliably in structured EHR data and which depended on missing, inconsistent, or institution-specific information.
Alternative evidence pathways were documented when the same clinical state could be represented through different combinations of events.
3. Concept-Set Curation
Candidate concepts were identified using clinical terminology knowledge, OMOP vocabulary relationships, lexical searches, mapping review, and LLM-assisted synonym expansion, then reviewed against the same criteria described under OMOP Concept-Set Engineering. Inclusion rationale was documented before the set was approved for implementation.
4. Cohort Implementation
Approved clinical logic was implemented as OMOP cohort definitions. We configured initial events, restrictions, inclusion rules, temporal relationships, observation requirements, exit criteria, and cohort-era behavior.
Derived clinical states that required post-cohort calculations were retained as downstream analytical logic rather than artificially represented as independent entry-event cohorts.
5. Technical and Semantic QA
We compared the original clinical specification with:
This review identified semantic omissions, implementation inconsistencies, incorrect temporal relationships, domain mismatches, duplicate-event behavior, and dependencies on unavailable data.
6. Cross-Site Evaluation
Aggregate outputs were compared across participating data sources. The analysis focused on:
Unexpected variation triggered focused review of data availability, mappings, and implementation assumptions.
7. Multidisciplinary Review
Clinical, vocabulary, observational, and engineering review ran in parallel, each catching a different category of failure. LLM support fed into this process without an approval role of its own.

The client gained a repeatable lifecycle covering clinical specification, evidence review, terminology engineering, cohort implementation, diagnostics, expert approval, versioning, and release.
Clinical assumptions, terminology choices, temporal relationships, implementation decisions, and known limitations were documented explicitly instead of remaining embedded in individual queries.
The project demonstrated how LLMs could support phenotype knowledge engineering and QA without transferring clinical authority from accountable human experts to a generative model.