LLM Inference Latency Estimator

Estimate the time-to-first-token and per-token generation latency for large language models. Inputs include model FLOPs, GPU TFLOP/s, batch size, and memory bandwidth.

Networking
Algorithms
Binary & Number
Systems
Dev Metrics

IP Subnet Calculator

IP Address
CIDR Prefix
/
Network
192.168.1.0
Broadcast
192.168.1.255
Subnet Mask
255.255.255.0
First Host
192.168.1.1
Last Host
192.168.1.254
Usable Hosts
254
Binary breakdown:
IP: 11000000.10101000.00000001.00000000
Mask: 11111111.11111111.11111111.00000000
Net: 11000000.10101000.00000001.00000000
Advertisement

How to Use

Enter the values for the patient or scenario you are assessing. Estimate the time-to-first-token and per-token generation latency for large language models. Inputs include model FLOPs, GPU TFLOP/s, batch size, and memory bandwidth. Use the inference latency result to inform your clinical assessment.

Frequently Asked Questions