Estimate end-to-end latency for a Retrieval-Augmented Generation (RAG) pipeline. Model query embedding, ANN retrieval, reranking, context assembly, and LLM generation latency.
Enter the values for the patient or scenario you are assessing. Estimate end-to-end latency for a Retrieval-Augmented Generation (RAG) pipeline. Model query embedding, ANN retrieval, reranking, context assembly, and LLM generation latency. Use the rag latency result to inform your clinical assessment.