ollama

mirror of https://github.com/ollama/ollama.git synced 2026-04-18 12:54:12 +02:00

Author	SHA1	Message	Date
Daniel Hiltgen	e38b606e8b	bench: add prompt calibration, context size flag, and NumCtx reporting Add --num-ctx flag to set context size, and report NumCtx in model info header. Calibrate tokens-per-word ratio during warmup using actual tokenization metrics from the model, replacing the fixed 1.3 heuristic. This produces more accurate prompt token counts for --prompt-tokens. Also add fetchContextLength() to query running model context via /api/ps.	2026-04-01 15:20:37 -07:00
Daniel Hiltgen	79c1e93c00	bench: improve benchmarking tool (#14240 ) New features: - Warmup phase to eliminate cold-start outliers - time-to-first-token measured in each epoch - VRAM/memory tracking to identify CPU spillover - Controlled prompt length - Defaults to 6 epochs and 200 tokens max Benchstat fixes: - ns/request instead of ns/op — non-standard unit created a separate group instead of grouping with timing metrics - Token count as the N field — benchstat interprets N as iteration count for statistical weighting, not as a token count	2026-03-15 11:47:31 -07:00
Eloi Torrents	a03223b86f	cmd/bench: support writing benchmark output to file (#13263 ) * cmd/bench: support writing benchmark output to file This changes Ollama to allow the bench command to write benchmark results to a user-specified output file instead of stdout when the --output flag is provided. --------- Co-authored-by: Patrick Devine <patrick@infrahq.com>	2025-12-04 13:22:41 -08:00
Patrick Devine	d7fd72193f	tests: basic benchmarking test framework (#12964 ) This change adds a basic benchmarking test framework for Ollama which can be used to determine the prefill, eval, load duration, and total duration for running a given model or models.	2025-11-15 18:17:40 -08:00

4 Commits