Calculate the inference latency overhead introduced by adapter layers and LoRA modules during deployment. Model the impact of adapter rank, number of adapted layers, and hardware characteristics on per-token generation speed.
Enter the values for the patient or scenario you are assessing. Calculate the inference latency overhead introduced by adapter layers and LoRA modules during deployment. Model the impact of adapter rank, number of adapted layers, and hardware characteristics on per-token generation speed. Use the adapter latency overhead result to inform your clinical assessment.