ollama

mirror of https://github.com/ollama/ollama.git synced 2026-04-17 21:54:08 +02:00

Author	SHA1	Message	Date
Devon Rifkin	e585ecd11f	gemma4: render differently based on model size Following up on #15560, this change now has e2b/e4b render differently from 26b/31b. For backwards compatibility, we take the existing renderer name `gemma4` and make it do dynamic resolution based on the model name/size, but the intended use is for the models to be republished with the renderer variant specified explicitly: `gemma4-small` or `gemma4-large`.	2026-04-15 14:37:16 -07:00
Daniel Hiltgen	06ae6367bd	mlx: fix RotatingKVCache.concat() dropping context on mid-rotation (#15591 ) After the rotating buffer has wrapped (c.offset > c.maxSize) a subsequent L>1 Update() went through a slice-to-[0, c.idx) path that discarded all slots in [c.idx, Dim), losing the older-but-still-in-window tokens the first Q of the new batch needs for its sliding-window attention. Linearize the circular buffer to logical order in that wrapped case so the existing trim + concat preserves the last (maxSize - 1) old tokens. When the buffer has not yet wrapped (c.offset <= c.maxSize), slots [c.idx, Dim) are grow padding or stale post-rewind data, so keep dropping them.	2026-04-14 18:29:06 -07:00
Daniel Hiltgen	48ad7085c4	mlx: Improve gemma4 performance with fused operations (#15587 ) * mlx: Improve gemma4 performance with fused operations * review comments	2026-04-14 18:04:04 -07:00
Jesse Gross	e1e3cec8d0	models: fuse MLP activation functions via mlx_compile Converts SiLU/GELUApprox to compiled kernels and adds SwiGLU, matching upstream mlx/mlx_lm's activations pattern. Routes llama, qwen3, qwen3_5 (dense + MoE), and glm4_moe_lite MLP paths through mlx.SwiGLU so each MLP invocation runs as one fused Metal/CUDA kernel rather than a chain of per-op launches.	2026-04-14 16:38:32 -07:00
Jesse Gross	d3e67e305c	mlx: add compiled closure support Wraps MLX's mlx_compile API so Go functions can be traced into fused kernels. Contiguous elementwise chains collapse into a single Metal/CUDA kernel instead of launching one per op. Exposes Compile plus arity helpers (Compile1/2/3) that mirror Python's @mx.compile decorator shape, lazily building the closure on first call so package-level declarations work before the MLX dylib loads.	2026-04-14 16:38:32 -07:00
Eva H	698e04a14b	launch: OpenCode inline config (#15586 )	2026-04-14 15:08:42 -07:00
Eva H	1d9537bc33	launch/openclaw: fix --yes flag behaviour to skip channels configuration (#15589 )	2026-04-14 13:57:35 -07:00
Eva H	120424d832	Revert "launch/opencode: use inline config (#15462 )" (#15568 )	2026-04-13 18:40:17 -07:00
Eva H	5818001610	launch: skip unchanged integration rewrite configration (#15491 )	2026-04-13 17:18:56 -07:00
Daniel Hiltgen	2cba7756c5	Gemma4 on MLX (#15244 ) * gemma4: implement Gemma 4 model for MLX (text-only runtime) * gemma4: two MoE + SWA prefill perf fixes Two performance optimizations in the gemma4 forward pass 1. Memoize the sliding-window prefill mask across layers. 2. Softmax only over the selected experts in Router.Forward. * review comments v0.20.8-rc0	2026-04-13 16:36:51 -07:00
Devon Rifkin	bf2a421727	gemma4: restore e2b-style nothink prompt (#15560 ) Gemma 4 prompts differ when thinking is disabled for different sized models: 26b/31b emit an empty thought block, while e2b/e4b do not. Before #15490, our shared Gemma 4 renderer effectively matched the e2b behavior. #15490 changed it to always emit the empty thought block, which regressed e2b/e4b nothink behavior and led to #15536 (and possibly This change restores the previous shared behavior by removing the empty trailing thought block. It also renames the checked-in upstream chat templates so the e2b and 31b fixtures are tracked separately. A follow-up will split Gemma 4 rendering by model size. Fixes: #15536	2026-04-13 14:26:15 -07:00
Eva H	f3cf6b75fb	launch/opencode: use inline config (#15462 )	2026-04-13 13:41:31 -07:00
Devon Rifkin	5dfac387a6	Revert "gemma4: fix nothink case renderer (#15553 )" (#15556 ) This reverts commit `4d75f5da03`.	2026-04-13 13:12:18 -07:00
Daniel Hiltgen	a99e5d9c22	mac: prevent generate on cross-compiles (#15120 ) For some versions of Xcode, cmake builds are failing due to header problems in cross-compiling during the generate phase. Since generate is producing arch independent generated output, we can skip this during cross-compiling.	2026-04-13 13:04:58 -07:00
Daniel Hiltgen	0abf3aca36	cgo: suppress deprecated warning to quiet down go build (#15438 )	2026-04-13 13:04:11 -07:00
Devon Rifkin	ee0266462a	Revert "gemma4: add nothink renderer tests (#15554 )" (#15555 ) This reverts commit `1b70bb8a10`.	2026-04-13 13:00:59 -07:00
Daniel Hiltgen	c88fb286ec	mlx: add op wrappers for Conv2d, Pad, activations, trig, and masked SDPA (#14913 ) * mlx: add op wrappers for Conv2d, Pad, activations, trig, and masked SDPA Add Conv2d, flexible Pad (with axes/mode), PadConstant, Maximum, Minimum, Softplus, ReLU, GLU, Clamp, Sin, Cos, Clip, ScaledDotProductAttentionMasked, and RoPEWithFreqs. Refactor RoPEWithBase to delegate to RoPEWithFreqs. * review comments * mlx: fix ScaledDotProductAttentionMasked to consult the mask argument	2026-04-13 11:43:24 -07:00
Daniel Hiltgen	d3da29cbfc	mlx: mixed-precision quant and capability detection improvements (#15409 ) Improve the MLX model creation pipeline with several model-agnostic changes: - Rewrite supportsVision to use vision_config instead of architecture name - Add supportsAudio for audio encoder detection - Add alignment checking (isAligned) for quantization group sizes - Support per-projection mixed quantization in MoE expert packing - Record per-tensor quant metadata in safetensors blobs - Parse per-tensor quant metadata at model load time - Validate quantize output is non-empty before storing - Fix pin/unpin cleanup in expert group quantization - Promote v_proj/k_proj/down_proj to INT8 for INT4 base quant - Add MetalIsAvailable() utility - Skip audio encoder tensors from quantization	2026-04-13 11:43:07 -07:00
Devon Rifkin	1b70bb8a10	gemma4: add nothink renderer tests (#15554 ) Meant to include in #15553 v0.20.7-rc0	2026-04-13 11:38:19 -07:00
Daniel Hiltgen	ec29ce4ce3	gemma4: fix compiler error on metal (#15550 ) On some systems, the metal runtime compiler is failing due to an uninitialized variable from #15378. Fixes #15548	2026-04-13 11:32:00 -07:00
Devon Rifkin	4d75f5da03	gemma4: fix nothink case renderer (#15553 ) Regressed in #15490 Fixes: #15536	2026-04-13 11:23:19 -07:00
saman-amd	798fd09bfe	Update to ROCm 7.2.1 (#15483 ) Co-authored-by: Samiii777 <58442200+Samiii777@users.noreply.github.com>	2026-04-12 12:11:58 -07:00
Devon Rifkin	9330bb9120	gemma4: be less strict about whitespace before bare keys (#15494 ) v0.20.6-rc1 v0.20.6	2026-04-11 16:30:27 -07:00
Devon Rifkin	40a1317dfd	gemma4: update renderer to match new jinja template (#15490 ) * gemma4: update renderer to match new jinja template Google has updated their jinja template for gemma4, and so this change gives us parity with the new template. The parsing also slightly changed upstream, so we make a small change to our parser as well. I've also corrected a few probably existing edge cases, especially around type unions. The upstream output format is weird (a stringified array), but in practice the models seem to understand it well. * gemma4: special case simple `AnyOf`s The upstream template doesn't handle `AnyOf`s, but since in the previous commit we saw type unions work reasonably well, I'm now treating very simple `AnyOf`s as type unions to help in cases where they might be used * fix lint * gemma4: prefer empty instead of `None` We can't currently distinguish between a result being not-present vs. empty. The empty case seems more important (e.g., a legitimately empty tool call) * gemma4: be more careful for tool results with missing IDs v0.20.6-rc0	2026-04-10 15:45:27 -07:00
Devon Rifkin	fdfe9cec98	model/parsers: fix missing parallel tool call indices (#15467 ) We were missing setting the function index for several models that can make parallel tool calls. In the future we may want to consider putting some sort of post-parse hook and relieve the parsers of this duty. Fixes: #15457	2026-04-10 15:23:21 -07:00
Matteo Celani	9517864603	app/ui: re-validate image attachments when selected model changes (#15272 )	2026-04-10 14:03:51 -07:00
Bruce MacDonald	8e6d86dbe3	docs: add hermes agent integration guide (#15488 ) Update cloud and local model recommendations to match current models.go: add qwen3.5:cloud and glm-5.1:cloud, replace glm-4.7-flash with gemma4 and qwen3.5 as local options. Add documentation for Hermes Agent by Nous Research, covering installation, Ollama setup via custom endpoint, messaging configuration, and recommended models.	2026-04-10 13:13:36 -07:00
Parth Sareen	80d3744c5d	launch: update openclaw channel message (#15463 ) v0.20.5-rc2 v0.20.5	2026-04-09 15:20:30 -07:00
Eva H	2a94f03823	launch: add re-run hint to dependency error message (#15439 ) v0.20.5-rc1	2026-04-09 09:51:34 -07:00
Patrick Devine	eb97274e5c	modelfiles: fix /save command and add shortname for safetensors based models (#15413 ) This change fixes two issues with Modelfiles: 1. If a user uses `ollama show --modelfile` to show a safetensors based model, the Model would leave the "FROM" field blank which won't allow a user to recreate the model. This change adds the model's current canonical short name to the FROM field. 2. If a user uses the `/save` command in the CLI any messages which were saved in a previous model wouldn't get saved (only the set of messages from the current session).	2026-04-08 21:05:39 -07:00
Daniel Hiltgen	6b5db12aa2	mlx: remove stale x86 libmlx library (#15443 ) Fixes #15433	2026-04-08 20:51:47 -07:00
7. Sun	612f0a17d3	fix: improve error message for unknown input item type in responses API (#15424 ) The default branch in unmarshalResponsesInputItem had two issues: - It referenced typeField.Type instead of itemType; these differ when the shorthand role-based format promotes an empty type to "message", meaning an unhandled type would show the wrong value in the error string. - It used %s formatting, so an empty type field produced the unhelpful message "unknown input item type: " with no indication what was missing. Fix by using itemType (the resolved value) with %q quoting, and add a dedicated message when itemType is empty (both type and role absent): "input item missing required 'type' field". Tests added for the empty-type and missing-type cases. Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> v0.20.5-rc0	2026-04-08 17:41:12 -07:00
Parth Sareen	673726fa0e	app: restore launch default and refine launch sidebar open for app (#15437 )	2026-04-08 16:59:21 -07:00
Daniel Hiltgen	b5918f9785	pull/push: refine safetensors (#14946 ) * pull: refine safetensors pull - Body drain in resolve() — drain response body before close so Go's HTTP client can reuse TCP connections instead of opening a new one per blob (1,075 extra TCP+TLS handshakes eliminated) - Skip speed recording for tiny blobs (<100KB) — prevents HTTP-overhead-dominated transfer times from poisoning the median, which the stall detector uses to cancel "too slow" downloads - Resume support for large blobs (>=64MB) — on failure, preserves partial .tmp files; on retry, re-hashes existing datak and sends Range header to download only remaining bytes; gracefully falls back to full download if server returns 200 instead of 206; SHA256 verification catches corrupt partials * harden push - Prevents killing TCP connections after every request. - Stronger backoff to handle server back-pressure and rate limiting - Larger buffered reads for improve safetensor upload performance - Better error message handling from server - Handle 201 if server says blob exists - Fix progress reporting on already uploaded blobs - Trace logging to help troubleshoot and tune going forward * review comments * review comments	2026-04-08 14:15:39 -07:00
Eva H	d17f482d50	launch/opencode: detect `curl` installed opencode at `~/.opencode/bin` (#15197 )	2026-04-08 13:54:51 -07:00
Parth Sareen	4e16f562c0	launch: add openclaw channels setup (#15407 )	2026-04-08 13:25:27 -07:00
Parth Sareen	55308f1421	launch: update ctx length for glm-5.1 and gemma4 (#15411 ) Also adds glm-5.1 in recommended models	2026-04-08 12:11:50 -07:00
Eva H	d64812eb5d	cmd: improve multi-select sorting and selection status (#15200 )	2026-04-08 10:39:18 -07:00
Devon Rifkin	f86a969f27	responses: add support for fn call output arrays (#15406 ) In addition to strings (which we already supported), OpenResponses supports arrays of text content, image content, or file content (see <https://www.openresponses.org/reference#object-FunctionCallOutput-title>). We were missing support for these arrays, which caused unmarshal errors like ``` json: cannot unmarshal array into Go struct field ResponsesFunctionCallOutput.output of type string ``` This change adds support for text content and image content, as those are more straightforwardly mappable to Ollama message formats (though image and text interleaving is lost), but it's less clear what to do for files. In the future we can partially support this by inlining reasonably sized text files, but wanted to get this change out first. Fixes: #15250 v0.20.4	2026-04-07 16:47:30 -07:00
Matteo Celani	9fa80a1660	app/ui: fix lint errors for unused vars, prefer-const, and empty catch (#15282 )	2026-04-07 16:28:36 -07:00
Daniel Hiltgen	dde09129d1	gemma4: Disable FA on older GPUs where it doesn't work (#15403 ) CUDA older than 7.5 lack the support to enable flash attention for the model. v0.20.4-rc2	2026-04-07 14:54:25 -07:00
Patrick Devine	780556c4d0	mlx: use default http client (#15405 )	2026-04-07 14:53:23 -07:00
Daniel Hiltgen	dfae363b5b	gemma4: add missing file (#15394 ) File accidentally omitted from #15378 v0.20.4-rc1	2026-04-07 09:18:01 -07:00
Daniel Hiltgen	30fdd229a4	create: Clean up experimental paths, fix create from existing safetensor model (#14679 ) * create: Clean up experimental paths This cleans up the experimental features, and adds both unit and integration test coverage to verify no regressions. * create: preserve config and layer names when creating from safetensors models When creating a model FROM an existing safetensors model, ModelFormat, Capabilities, and layer Name fields were lost. ModelFormat stayed empty because it's only set from GGML layers (which safetensors models lack), and layer names weren't copied in parseFromModel. This caused derived models to fail loading ("config.json not found in manifest"). * review comments v0.20.4-rc0	2026-04-07 08:12:57 -07:00
Daniel Hiltgen	e823bff873	gemma4: enable flash attention (#15378 ) Backport GGML kernels so we can enable flash attention for the gemma 4 model on Metal and CUDA.	2026-04-07 08:12:36 -07:00
Daniel Hiltgen	8968740836	mlx: Improve M5 performance with NAX (#15345 ) * mlx: Improve M5 performance with NAX This modifies the Mac release to now have 2 builds of MLX for broader compatibility while supporting the latest M5 hardware features. NAX requires building with xcode 26.2 and targetting support only for OS v26 and up. Since we want to support older MacOS versions as well, we now need 2 different MLX builds and runtime detection logic to select the optimal version. The newer build will detect NAX missing at runtime, so it is safe to run on pre M5 macs. * mac: prevent generate on cross-compiles For some versions of Xcode, cmake builds are failing due to header problems in cross-compiling during the generate phase. Since generate is producing arch independent generated output, we can skip this during cross-compiling.	2026-04-07 08:12:24 -07:00
Devon Rifkin	8c8f8f3450	model/parsers: add gemma4 tool call repair (#15374 ) The existing strict gemma4 tool parser is still the primary path, but if this fails, we try to repair by fixing some of the most commonly seen mistakes these models seem to make in practice. We repair by building up a set of candidates, and use the first candidate that parses. Repairs cover: - missing Gemma string delimiters - single-quoted string values, including a dangling Gemma delimiter - raw terminal string values (if the corresponding tool schema indicates it should be a string) - missing object close only after a concrete repair Add regression coverage for malformed tool calls from issue #15315 and focused unit tests for the individual repair helpers and candidate pipeline. v0.20.3-rc0 v0.20.3	2026-04-06 18:47:17 -07:00
Parth Sareen	82f0139587	launch/openclaw: patch approvedScopes baseline for TUI pairing (#15375 )	2026-04-06 18:00:12 -07:00
Bruce MacDonald	26a58b294c	app: update featured models (#15373 ) Featured models in the app are out of date. Update them to a more recent list of models.	2026-04-06 16:35:35 -07:00
Devon Rifkin	34a790a2e6	model/parsers: suppress extra gemma4 closing tool tags (#15370 ) We've observed Gemma 4 occasionally emitting extra <tool_call\|> tags after a valid tool call. We suppress leading close tags in this immediate post-tool-call state so the extra close tags do not leak into assistant content. The tradeoff is that if the model intentionally begins its next content span with the literal string "<tool_call\|>", we will erroneously treat it as noise and drop it.	2026-04-06 12:41:33 -07:00

1 2 3 4 5 ...

5312 Commits