ollama

mirror of https://github.com/ollama/ollama.git synced 2026-04-25 18:25:42 +02:00

Author	SHA1	Message	Date
jmorganca	61b367ec29	llama/compat: shrink patch to pure call-site hooks (34 -> 20 lines) Two reductions: 1. Drop the gguf_rename_tensor forwarder from gguf.h/gguf.cpp. The rename-in-place trick it does (calling ggml_set_name on an embedded ggml_tensor) can be done from outside gguf.cpp via: char * p = const_cast<char *>(gguf_get_tensor_name(meta, id)); strncpy(p, new_name, GGML_MAX_NAME - 1); That pointer aims into a mutable char[GGML_MAX_NAME] inside a std::vector element; the const on the return type is API courtesy. This is defined behavior and has no struct-layout dependency. 2. Drop the src/CMakeLists.txt hunk that added llama-ollama-compat.cpp to the llama target. Replace with a target_sources() call in Ollama's llama/server/CMakeLists.txt after FetchContent_MakeAvailable. Our compat files now stay in llama/compat/ and are never copied into the fetched _deps/ tree. Net patch now touches 3 files, 20 lines, all pure call-site insertions: src/llama-model-loader.cpp +8 (include + translate + 2x should_skip) src/llama-model.cpp +4 (include + apply_tensor_transforms) tools/mtmd/clip.cpp +8 (include + translate_clip + maybe_load) Verified: fresh build from scratch (rm -rf build && cmake configure) runs PATCH_COMMAND cleanly, compiles, and ollama run gemma3 still works end-to-end for text + vision.	2026-04-20 09:29:34 -07:00
jmorganca	021389f7bb	llama/compat: shrink clip.cpp injection from 18 lines to 1 The clip.cpp tensor-read loop was the fattest hook in the patch — it duplicated the host-vs-device buffer dispatch around a call into the compat layer. Move that dispatch into our code (maybe_load_tensor), so the upstream patch is a single conditional call. Net: upstream patch drops from 48 lines across 6 files to 34 lines. Every remaining edit is either a 1-line include, a 1-line function call, or the gguf_rename_tensor shim (which accesses gguf_context internals and has to live in gguf.cpp). Verified end-to-end: text + vision both still correct after rebuild.	2026-04-20 09:29:34 -07:00
jmorganca	8c2c9d4c89	llama/compat: extend gemma3 handler to cover 1B and 270M blobs Previous handler only fired on vision-capable gemma3 (4B/12B/27B) because its detection looked for `gemma3.mm.tokens_per_image` or embedded v./mm. tensors. The 1B blob has neither — but its old Ollama converter emitted: - gemma3.rope.global.freq_base (upstream uses gemma3.rope.freq_base) - gemma3.rope.local.freq_base (upstream uses gemma3.rope.freq_base_swa) - tokenizer.ggml.add_{padding,unknown}_token so llama.cpp would fall back to default rope_freq_base=10000 and produce visibly-worse output. Also inject rope.scaling.factor=8.0 / type=linear on 4B/12B/27B — those variants ship with that scaling in their HF config to extend the native ~16k trained context to 131072. Without this KV, llama.cpp uses factor=1.0 and the positional embeddings are subtly off everywhere. Detection now flips on any Ollama-specific marker. All three variants verified end-to-end via `ollama run gemma3:{latest,1b,270m}`.	2026-04-20 09:29:34 -07:00
jmorganca	436f2e2b15	llama/compat: make patch-apply idempotent FetchContent's PATCH_COMMAND runs after each update step — including on incremental rebuilds. `git apply` fails when the patch is already applied, which bricks the build until the dev wipes build/ entirely. Fix by routing the apply through a small apply-patch.cmake helper that checks `git apply --reverse --check` first. If the patch cleanly reverses, it's already applied and we skip. Otherwise apply forward. Both branches surface real errors (drift against upstream, missing patch file, etc.). Verified: fresh configure+build applies the patch once; re-running the same commands is a no-op with no errors.	2026-04-20 09:29:34 -07:00
jmorganca	7449b539ab	llm,server: route Ollama-format gemma3 blobs through llama/compat Two tiny Go-side changes that let the llama/compat shim take over gemma3: 1. llm/llama_server.go: when the GGUF has embedded v.* tensors and no projector layer is declared, pass the model file itself as --mmproj. The in-process compat layer translates the same file into both a text-only view (for --model) and a clip-mmproj view (for --mmproj). 2. server/model_resolver.go: drop library/gemma3 from compatModelRedirects. The compat layer handles it directly, so no dhiltgen/ republish is needed. Other arches stay in the redirect list until they get their own handler in llama/compat/llama-ollama-compat.cpp. End-to-end verified: `ollama run gemma3` answers text and image prompts against the existing library/gemma3 blob with no re-download.	2026-04-20 09:29:34 -07:00
jmorganca	25223160d8	llama/compat: add in-memory shim so llama-server can load Ollama-format GGUFs Older Ollama builds ship GGUFs that diverge slightly from upstream llama.cpp in arch names, KV keys, tensor names, and (for vision models) file layout (text+vision in one monolithic file). This adds a self-contained compat layer that translates those files in memory at load time, so ~/.ollama/models/blobs/* can be served by upstream llama-server with no re-conversion and no re-download. Structure: llama/compat/ llama-ollama-compat.{h,cpp} — the shim (Ollama-owned, ~500 LOC) upstream-edits.patch — ~48 lines of call-site hooks in 6 upstream files compat.cmake — include()-able CMake fragment README.md — what/why/how-to-regen Integration: llama/server/CMakeLists.txt includes compat.cmake and passes OLLAMA_LLAMA_CPP_COMPAT_PATCH_COMMAND to FetchContent_Declare via PATCH_COMMAND. When OLLAMA_LLAMA_CPP_SOURCE is set (dev mode), the patch is skipped so the developer's tree stays untouched. Currently handles gemma3 (text + vision). Pattern is data-driven — adding other archs is a new handle_<arch>() + one dispatch line. See README for the per-arch checklist. Verified end-to-end: `llama-server --model BLOB --mmproj BLOB` with an Ollama gemma3:latest blob answers both text prompts ("Paris") and vision prompts (correct image descriptions).	2026-04-20 09:29:34 -07:00
Daniel Hiltgen	56c735d871	runner: Remove CGO engines, use llama-server exclusively for GGML models Remove the vendored GGML and llama.cpp backend, CGO runner, Go model implementations, and sample. llama-server (built from upstream llama.cpp via FetchContent) is now the sole inference engine for GGUF-based models. (Safetensor based models continue to run on the new MLX engine.) This allows us to more rapidly pick up new capabilities and fixes from llama.cpp as they come out. On windows this now requires recent AMD driver versions to support ROCm v7 as llama.cpp currently does not support building against v6.	2026-04-20 08:44:02 -07:00
Daniel Hiltgen	ff23dd343f	mlx: apply repeat penalties in sampler (#15631 )	2026-04-18 07:49:38 -07:00
Parth Sareen	123b300af6	docs: update hermes (#15655 )	2026-04-17 14:20:59 -07:00
Parth Sareen	57653b8e42	cmd/launch: show WSL guidance on Windows instead of handing off (#15637 ) v0.21.0-rc1 v0.21.0	2026-04-16 17:18:04 -07:00
Parth Sareen	a50ce61c54	launch: skip unchanged managed-single rewrite (#15633 )	2026-04-16 16:20:42 -07:00
Daniel Hiltgen	2bb7ea00d2	create: avoid gc race with create (#15628 ) If you have a long running create, and start another ollama server with the same model dir, the GC algorithm deletes the pending blobs and breaks the create. This adds a 1h grace period to avoid deleting in-flight creation operations.	2026-04-16 13:29:16 -07:00
Daniel Hiltgen	55fa80d07a	mlx: additional gemma4 cache fixes (#15607 ) Harden additional corner cases	2026-04-16 13:07:19 -07:00
Daniel Hiltgen	b9cb535407	mlx: fix gemma4 cache to use logical view (#15617 ) v0.21.0-rc0	2026-04-16 11:54:30 -07:00
Daniel Hiltgen	031baef094	mlx: fix imagegen lookup (#15588 ) * mlx: fix imagegen lookup Fixes #15533 - imagegen had fallen out of sync with the new layout for multiple mlx libraries on Metal. * review comments	2026-04-16 10:39:00 -07:00
Mike Wallio	7d271e6dc9	cmd/launch: add Copilot CLI integration (#15583 ) --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: ParthSareen <parth.sareen@ollama.com>	2026-04-15 17:22:53 -07:00
Devon Rifkin	c88dae2d6b	Merge pull request #15612 from ollama/drifkin/gemma4-split-templates gemma4: render differently based on model size	2026-04-15 17:15:35 -07:00
Devon Rifkin	9e3618d663	make empty block conditional	2026-04-15 15:35:25 -07:00
Daniel Hiltgen	5d920cc6bc	Keep Gemma4 router projection in source precision (#15613 )	2026-04-15 15:04:23 -07:00
Devon Rifkin	e585ecd11f	gemma4: render differently based on model size Following up on #15560, this change now has e2b/e4b render differently from 26b/31b. For backwards compatibility, we take the existing renderer name `gemma4` and make it do dynamic resolution based on the model name/size, but the intended use is for the models to be republished with the renderer variant specified explicitly: `gemma4-small` or `gemma4-large`.	2026-04-15 14:37:16 -07:00
Eva H	cdddea0592	launch: always list cloud recommendations first (#15593 )	2026-04-15 13:17:35 -07:00
Parth Sareen	43f90def04	launch: add hermes (#15569 )	2026-04-15 12:00:23 -07:00
Daniel Hiltgen	06ae6367bd	mlx: fix RotatingKVCache.concat() dropping context on mid-rotation (#15591 ) After the rotating buffer has wrapped (c.offset > c.maxSize) a subsequent L>1 Update() went through a slice-to-[0, c.idx) path that discarded all slots in [c.idx, Dim), losing the older-but-still-in-window tokens the first Q of the new batch needs for its sliding-window attention. Linearize the circular buffer to logical order in that wrapped case so the existing trim + concat preserves the last (maxSize - 1) old tokens. When the buffer has not yet wrapped (c.offset <= c.maxSize), slots [c.idx, Dim) are grow padding or stale post-rewind data, so keep dropping them.	2026-04-14 18:29:06 -07:00
Daniel Hiltgen	48ad7085c4	mlx: Improve gemma4 performance with fused operations (#15587 ) * mlx: Improve gemma4 performance with fused operations * review comments	2026-04-14 18:04:04 -07:00
Jesse Gross	e1e3cec8d0	models: fuse MLP activation functions via mlx_compile Converts SiLU/GELUApprox to compiled kernels and adds SwiGLU, matching upstream mlx/mlx_lm's activations pattern. Routes llama, qwen3, qwen3_5 (dense + MoE), and glm4_moe_lite MLP paths through mlx.SwiGLU so each MLP invocation runs as one fused Metal/CUDA kernel rather than a chain of per-op launches.	2026-04-14 16:38:32 -07:00
Jesse Gross	d3e67e305c	mlx: add compiled closure support Wraps MLX's mlx_compile API so Go functions can be traced into fused kernels. Contiguous elementwise chains collapse into a single Metal/CUDA kernel instead of launching one per op. Exposes Compile plus arity helpers (Compile1/2/3) that mirror Python's @mx.compile decorator shape, lazily building the closure on first call so package-level declarations work before the MLX dylib loads.	2026-04-14 16:38:32 -07:00
Eva H	698e04a14b	launch: OpenCode inline config (#15586 )	2026-04-14 15:08:42 -07:00
Eva H	1d9537bc33	launch/openclaw: fix --yes flag behaviour to skip channels configuration (#15589 )	2026-04-14 13:57:35 -07:00
Eva H	120424d832	Revert "launch/opencode: use inline config (#15462 )" (#15568 )	2026-04-13 18:40:17 -07:00
Eva H	5818001610	launch: skip unchanged integration rewrite configration (#15491 )	2026-04-13 17:18:56 -07:00
Daniel Hiltgen	2cba7756c5	Gemma4 on MLX (#15244 ) * gemma4: implement Gemma 4 model for MLX (text-only runtime) * gemma4: two MoE + SWA prefill perf fixes Two performance optimizations in the gemma4 forward pass 1. Memoize the sliding-window prefill mask across layers. 2. Softmax only over the selected experts in Router.Forward. * review comments v0.20.8-rc0	2026-04-13 16:36:51 -07:00
Devon Rifkin	bf2a421727	gemma4: restore e2b-style nothink prompt (#15560 ) Gemma 4 prompts differ when thinking is disabled for different sized models: 26b/31b emit an empty thought block, while e2b/e4b do not. Before #15490, our shared Gemma 4 renderer effectively matched the e2b behavior. #15490 changed it to always emit the empty thought block, which regressed e2b/e4b nothink behavior and led to #15536 (and possibly This change restores the previous shared behavior by removing the empty trailing thought block. It also renames the checked-in upstream chat templates so the e2b and 31b fixtures are tracked separately. A follow-up will split Gemma 4 rendering by model size. Fixes: #15536	2026-04-13 14:26:15 -07:00
Eva H	f3cf6b75fb	launch/opencode: use inline config (#15462 )	2026-04-13 13:41:31 -07:00
Devon Rifkin	5dfac387a6	Revert "gemma4: fix nothink case renderer (#15553 )" (#15556 ) This reverts commit `4d75f5da03`.	2026-04-13 13:12:18 -07:00
Daniel Hiltgen	a99e5d9c22	mac: prevent generate on cross-compiles (#15120 ) For some versions of Xcode, cmake builds are failing due to header problems in cross-compiling during the generate phase. Since generate is producing arch independent generated output, we can skip this during cross-compiling.	2026-04-13 13:04:58 -07:00
Daniel Hiltgen	0abf3aca36	cgo: suppress deprecated warning to quiet down go build (#15438 )	2026-04-13 13:04:11 -07:00
Devon Rifkin	ee0266462a	Revert "gemma4: add nothink renderer tests (#15554 )" (#15555 ) This reverts commit `1b70bb8a10`.	2026-04-13 13:00:59 -07:00
Daniel Hiltgen	c88fb286ec	mlx: add op wrappers for Conv2d, Pad, activations, trig, and masked SDPA (#14913 ) * mlx: add op wrappers for Conv2d, Pad, activations, trig, and masked SDPA Add Conv2d, flexible Pad (with axes/mode), PadConstant, Maximum, Minimum, Softplus, ReLU, GLU, Clamp, Sin, Cos, Clip, ScaledDotProductAttentionMasked, and RoPEWithFreqs. Refactor RoPEWithBase to delegate to RoPEWithFreqs. * review comments * mlx: fix ScaledDotProductAttentionMasked to consult the mask argument	2026-04-13 11:43:24 -07:00
Daniel Hiltgen	d3da29cbfc	mlx: mixed-precision quant and capability detection improvements (#15409 ) Improve the MLX model creation pipeline with several model-agnostic changes: - Rewrite supportsVision to use vision_config instead of architecture name - Add supportsAudio for audio encoder detection - Add alignment checking (isAligned) for quantization group sizes - Support per-projection mixed quantization in MoE expert packing - Record per-tensor quant metadata in safetensors blobs - Parse per-tensor quant metadata at model load time - Validate quantize output is non-empty before storing - Fix pin/unpin cleanup in expert group quantization - Promote v_proj/k_proj/down_proj to INT8 for INT4 base quant - Add MetalIsAvailable() utility - Skip audio encoder tensors from quantization	2026-04-13 11:43:07 -07:00
Devon Rifkin	1b70bb8a10	gemma4: add nothink renderer tests (#15554 ) Meant to include in #15553 v0.20.7-rc0	2026-04-13 11:38:19 -07:00
Daniel Hiltgen	ec29ce4ce3	gemma4: fix compiler error on metal (#15550 ) On some systems, the metal runtime compiler is failing due to an uninitialized variable from #15378. Fixes #15548	2026-04-13 11:32:00 -07:00
Devon Rifkin	4d75f5da03	gemma4: fix nothink case renderer (#15553 ) Regressed in #15490 Fixes: #15536	2026-04-13 11:23:19 -07:00
saman-amd	798fd09bfe	Update to ROCm 7.2.1 (#15483 ) Co-authored-by: Samiii777 <58442200+Samiii777@users.noreply.github.com>	2026-04-12 12:11:58 -07:00
Devon Rifkin	9330bb9120	gemma4: be less strict about whitespace before bare keys (#15494 ) v0.20.6-rc1 v0.20.6	2026-04-11 16:30:27 -07:00
Devon Rifkin	40a1317dfd	gemma4: update renderer to match new jinja template (#15490 ) * gemma4: update renderer to match new jinja template Google has updated their jinja template for gemma4, and so this change gives us parity with the new template. The parsing also slightly changed upstream, so we make a small change to our parser as well. I've also corrected a few probably existing edge cases, especially around type unions. The upstream output format is weird (a stringified array), but in practice the models seem to understand it well. * gemma4: special case simple `AnyOf`s The upstream template doesn't handle `AnyOf`s, but since in the previous commit we saw type unions work reasonably well, I'm now treating very simple `AnyOf`s as type unions to help in cases where they might be used * fix lint * gemma4: prefer empty instead of `None` We can't currently distinguish between a result being not-present vs. empty. The empty case seems more important (e.g., a legitimately empty tool call) * gemma4: be more careful for tool results with missing IDs v0.20.6-rc0	2026-04-10 15:45:27 -07:00
Devon Rifkin	fdfe9cec98	model/parsers: fix missing parallel tool call indices (#15467 ) We were missing setting the function index for several models that can make parallel tool calls. In the future we may want to consider putting some sort of post-parse hook and relieve the parsers of this duty. Fixes: #15457	2026-04-10 15:23:21 -07:00
Matteo Celani	9517864603	app/ui: re-validate image attachments when selected model changes (#15272 )	2026-04-10 14:03:51 -07:00
Bruce MacDonald	8e6d86dbe3	docs: add hermes agent integration guide (#15488 ) Update cloud and local model recommendations to match current models.go: add qwen3.5:cloud and glm-5.1:cloud, replace glm-4.7-flash with gemma4 and qwen3.5 as local options. Add documentation for Hermes Agent by Nous Research, covering installation, Ollama setup via custom endpoint, messaging configuration, and recommended models.	2026-04-10 13:13:36 -07:00
Parth Sareen	80d3744c5d	launch: update openclaw channel message (#15463 ) v0.20.5-rc2 v0.20.5	2026-04-09 15:20:30 -07:00
Eva H	2a94f03823	launch: add re-run hint to dependency error message (#15439 ) v0.20.5-rc1	2026-04-09 09:51:34 -07:00

1 2 3 4 5 ...

5333 Commits