Estimate the time-to-first-token and per-token generation latency for large language models. Inputs include model FLOPs, GPU TFLOP/s, batch size, and memory bandwidth.
Enter the values for the patient or scenario you are assessing. Estimate the time-to-first-token and per-token generation latency for large language models. Inputs include model FLOPs, GPU TFLOP/s, batch size, and memory bandwidth. Use the inference latency result to inform your clinical assessment.