Calculate the maximum request throughput for a deployed LLM serving system. Account for batch size, sequence length, GPU TFLOP/s, memory bandwidth, and continuous batching efficiency.
Enter the values for the patient or scenario you are assessing. Calculate the maximum request throughput for a deployed LLM serving system. Account for batch size, sequence length, GPU TFLOP/s, memory bandwidth, and continuous batching efficiency. Use the model serving throughput result to inform your clinical assessment.