Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization
Researchers present a benchmark for tokenization in generative medical event models, highlighting the importance of design choices in tokenization and event encoding for downstream tasks. The study evaluates the impact of quantization granularity, reference-range anchoring, and other tokenization methods on performance, finding that fused tokens and alternatives to explicit time tokens can improve model performance.
Save an API key to vote.