Search papers, labs, and topics across Lattice.
This paper introduces a hybrid neural-symbolic pipeline for extracting clinical follow-up instructions (action, date) from outpatient notes, addressing the limitations of generative models in handling temporal reasoning. The pipeline combines BioBERT-based entity extraction and linking with deterministic canonicalization and time normalization. Results on a synthetic dataset demonstrate that this hybrid approach significantly outperforms zero-shot GPT-4o-mini and fine-tuned LLaMA-3 8B in Test-Time Pair F1, particularly on out-of-vocabulary actions, while achieving near-perfect date accuracy.
Generative models stumble when extracting clinical follow-up instructions, but a hybrid neural-symbolic approach nails it, achieving near-perfect F1 and date accuracy even on unseen actions.
Objective. Outpatient notes carry follow-up instructions pairing actions with future times ("MRI brain in two weeks"). Extracting (action, date) pairs supports scheduling and audit, but generative extractors miss the date because linking and arithmetic are implicit in decoding. We test a hybrid neural-symbolic pipeline against direct generation. Methods. We define TestSpecification and TimeSpecification entities and a ScheduledFor relation. BioBERT feeds BIO tagging and a biaffine linker; entities are canonicalized via a 28-action ontology and times normalized to day offsets deterministically. We evaluate on a 2,000-note synthetic outpatient corpus with action-disjoint splits (18 train, 6 OOV-test) against zero-shot GPT-4o-mini and LoRA-fine-tuned LLaMA-3 8B with note-level bootstrap 95% CIs. Results. On 259-note seen and OOV splits the hybrid pipeline achieves Test-Time Pair F1 of 0.997 and 0.986 with 0.00-day MAE. Baselines reach high action F1 (LLaMA-3 0.992; GPT-4o-mini 0.963 seen) but Pair F1 stays at 0.51-0.57 (LLaMA-3) and 0.53 (GPT-4o-mini), CIs non-overlapping with the hybrid. Conclusion. Separating learned entity extraction from deterministic date arithmetic outperforms generation on this benchmark, generalizes to held-out actions, and exposes failure modes. Transfer to real EHR notes is the next validation; a first-pass realism check is in Limitations.