Calculate the token-per-second throughput of an LLM deployment. Enter GPU FLOP/s, memory bandwidth, model parameters, and serving batch size to benchmark expected performance.
Enter the values for the patient or scenario you are assessing. Calculate the token-per-second throughput of an LLM deployment. Enter GPU FLOP/s, memory bandwidth, model parameters, and serving batch size to benchmark expected performance. Use the tokens/s result to inform your clinical assessment.