Training GlucoFM to understand metabolism
We pre-trained GlucoFM on 109,066 hours of unlabeled CGM data from Wear-CGM
CGM recordings can contain gaps, different sampling intervals, and sensor artifacts. GlucoFM aligns each recording to a 24-hour, five-minute grid and retains an observation mask, keeping measured and unobserved positions distinct. Its dual-stream encoder separates a lower-frequency state component, representing slower glycemic trends, from a residual event component capturing short-term deviations that may arise from physiology, behavior, or sensing artifacts.
Rather than reconstructing exact raw glucose readings, which can be affected by measurement noise and sensor artifacts, GlucoFM uses latent predictive pre-training with two complementary tasks:
- Contextual prediction: We mask (i.e., hide) parts of a daily glucose sequence and ask the model to predict their latent representations from the surrounding context. By predicting in latent space, the model captures broader daily glucose patterns without having to reconstruct every sensor reading.
- Temporal dynamics: We also train the model to predict how a person’s steady baseline and short-term deviations will shift from one hour to the next. This encourages it to capture the continuous nature of glucose dynamics rather than treating readings as isolated snapshots in time.
Finally, CGM-aware augmentations introduce baseline drift, compression-like drops, sparser sampling, and short disconnections, exposing the model to variation and missingness encountered in real CGM recordings.
💸 Earn Instantly With This Task
No fees, no waiting — your earnings could be 1 click away.
Start Earning