The GRACE Cycle: A General Large-Language-Model Framework for Phenotype Discovery with Unknown Cluster Number.
The GRACE Cycle: A General Large-Language-Model Framework for Phenotype Discovery with Unknown Cluster Number.
Where did the research take place?
The study site has not been established. Author addresses may differ from where the research occurred.
Bethesda, US · Author affiliation
National Library of Medicine, Bethesda, MD, USA.Location evidence
Urbana, US · Author affiliation
University of Illinois Urbana-Champaign, Urbana, IL, USA.Location evidence
New York City, US · Author affiliation
Columbia University, New York, NY, USA.Location evidence
US · Author affiliation · country only
George Washington University, Washington, DC, USA.Location evidence
Hoboken, US · Author affiliation
Department of Computer Science, Stevens Institute of Technology, Hoboken, NJ, USA.Location evidence
Publication status: preprint
A plain-language reading has not been prepared for this paper yet.
Original abstract
Phenotype discovery-the data-driven identification of clinically or biologically meaningful subgroups-is fundamental to precision medicine, but conventional clustering methods require the number of clusters K to be specified a priori and struggle with heterogeneous, multimodal, or longitudinal data. We introduce the GRACE Cycle (Generate hypothesis, Retrieve evidence, Align, Converge, Evaluate), a general large-language-model (LLM)-assisted framework for phenotype discovery in which a hypothesis, an LLM, and an evidence base are iteratively refined until they agree. The framework discovers K as an output through Graph-of-Thought (GoT) refinement, in which an LLM reads per-cluster summary cards plus a between-cluster similarity matrix and proposes one of three moves-SPLIT, MERGE, or COMMIT-over a spectral-clustering seed. Two technical contributions enable scale: (i) a four-component prompt template integrating pairwise comparison, fairness pre-processing, and structured JSON output, and (ii) a data-feeding strategy that compresses cohorts of 10 4 - 10 6 entities into context-budget-respecting batches via k -nearest-neighbour graph sampling. We validate GRACE across three heterogeneous phenotyping problems: (1) longitudinal Long COVID subphenotyping in the NIH RECOVER cohort ( n = 13,511 ) , where GRACE recovers three clinically distinct subphenotypes (Protected, Responder, Refractory) with bootstrap Jaccard stability > 0.97 that are explained by a single autonomic/post-viral-fatigue axis (a 25-fold dysautonomia gradient, dysautonomia adjusted O R = 13.4 ) and an accompanying collapse of wearable-measured physical activity; (2) motor subphenotyping of Parkinson's disease from foot-sensor gait wearables (PhysioNet gaitpdb, n = 93 ), where GRACE discovers two gait subtypes without specifying K that are externally validated against withheld Timed-Up-and-Go ( p = 0.002 ) , Hoehn-Yahr stage ( p = 0.03 ) , and age; and (3) additional open wearable chronic-disease cohorts processed with the identical pipeline. Across domains, GRACE converges without prior knowledge of K , demonstrating that LLM-guided iterative reasoning offers a domain-agnostic alternative to conventional clustering when the number of phenotypes is unknown.