From cd86581a2a0a9ac08e86c8b4e565da3ee76569c7 Mon Sep 17 00:00:00 2001 From: Claude Date: Thu, 23 Jul 2026 09:39:53 +0000 Subject: [PATCH] Upgrade llama.cpp from b10076 to b10092 Bump the pinned llama.cpp release across the four canonical files (GIT_TAG + LLAMA_TAG in llama/CMakeLists.txt, README badge/link, CLAUDE.md pinned-version line, LlamaCppVersion.LLAMA_CPP_VERSION) and append the b10076->b10092 rows to the breaking-changes history. The b10076...b10092 range is 48 files / ~4.8k insertions over 16 commits, additive/tuning-only. All eight priority-8 headers are byte-identical. Two patch-target files changed but only outside the patched regions: common/arg.cpp adds speculative sidecar draft-repo resolution in common_models_handler_apply (away from 0001's common_params_parse* block; common/arg.h unchanged), and tools/server/server-context.cpp adds a null-ctx_tgt guard in load_model (distinct from 0002's callback guard and 0003's getters). All six patches (0001-0003, 0006-0008) apply cleanly in filename order against a fresh b10092 checkout, and the OuteTTS generator ran clean against tools/tts/tts.cpp @ b10092 (all anchors held). Co-Authored-By: Claude Opus 4.8 Claude-Session: https://claude.ai/code/session_01L7znBvN71TQNXgLjEGR9eU --- CLAUDE.md | 8 ++++---- README.md | 2 +- docs/history/llama-cpp-breaking-changes.md | 2 ++ llama/CMakeLists.txt | 4 ++-- .../java/net/ladenthin/llama/value/LlamaCppVersion.java | 8 ++++---- 5 files changed, 13 insertions(+), 11 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index fc3f4f635..a203058d2 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b10076** +Current llama.cpp pinned version: **b10092** ## Upgrading CUDA Version @@ -421,7 +421,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network; embed.cpp is plain C++17 (no npm) -git clone --depth 1 --branch b10076 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b10092 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build \ && ( cd dist && find . -type f -not -path './_gzip/*' \ | while read -r f; do mkdir -p "_gzip/$(dirname "$f")"; gzip -9 -c "$f" > "_gzip/$f"; done ) \ @@ -461,7 +461,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b10076`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b10092`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1254,7 +1254,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b10076`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b10092`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index ca70ee52b..2931b236b 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b10076](https://img.shields.io/badge/llama.cpp-%23b10076-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b10076) +[![llama.cpp b10092](https://img.shields.io/badge/llama.cpp-%23b10092-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b10092) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index b9d7da4e2..02da78be5 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -509,3 +509,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10069–b10075 | upstream verification (sandbox) | The `b10069...b10075` diff touches only the Hexagon backend, `src/llama-kv-cache-dsv4.cpp`, and `tools/ui/**` — none overlapping `0001`'s `common/arg.{cpp,h}` region, the `tools/server/*` targets of `0002`/`0003`/`0006`/`0007`/`0008`, or the OuteTTS generator anchors (`tools/tts/tts.cpp` unchanged), so all **six** patches (`0001`–`0003`, `0006`–`0008`) apply cleanly in filename order. Per-platform build + `ctest` confirmation by the CI pipeline. | | b10075–b10076 | `ggml/src/ggml-cuda/getrows.cu` | **Additive/perf-only, no public-API surface (1 file, 1 commit).** The sole changed file is a ggml-CUDA kernel: `k_get_rows_float` hoists loop-invariant index math out of the inner loop and a new `int4`-vectorized `k_get_rows_float_vec` path is taken in `get_rows_cuda_float` when alignment/size/block constraints hold (scalar kernel preserved as fallback). Inside an upstream-compiled CUDA backend TU — this project's CUDA classifiers are build-only, no source touched. All **eight** priority-8 headers (`common/{common.h,chat.h,speculative.h,arg.h,download.h}`, `tools/mtmd/mtmd.h`, `include/{llama.h,llama-cpp.h}`) are **byte-identical**; no `common/arg.*`, `tools/server/*`, or `tools/tts/tts.cpp` touched. All **six** patches (`0001`–`0003`, `0006`–`0008`) apply unchanged and the OuteTTS generator anchors hold. b10076 is the topmost release at bump time. | | b10075–b10076 | upstream verification (sandbox) | The `b10075...b10076` diff touches only `ggml/src/ggml-cuda/getrows.cu` — no overlap with `0001`'s `common/arg.{cpp,h}` region, the `tools/server/*` targets of `0002`/`0003`/`0006`/`0007`/`0008`, or the OuteTTS anchors (`tools/tts/tts.cpp` unchanged), so all **six** patches apply cleanly in filename order. Per-platform build + `ctest` confirmation by the CI pipeline. | +| b10076–b10092 | `common/arg.cpp` + `common/{chat-auto-parser-generator.cpp,chat-auto-parser.h,chat-diff-analyzer.cpp}` + `ggml/src/{ggml-cpu/{kleidiai/kleidiai.cpp,llamafile/sgemm.cpp},ggml-cuda/{common.cuh,convert.cu,dequantize.cuh,getrows.cu,ggml-cuda.cu,topk-moe.{cu,cuh}},ggml-hexagon/ggml-hexagon.cpp,ggml-openvino/ggml-openvino.cpp,ggml-vulkan/ggml-vulkan.cpp,ggml-webgpu/**}` + `src/{llama-arch.{cpp,h},llama-model.cpp,llama-model-saver.cpp,llama-vocab.{cpp,h},models/{laguna.cpp,models.h}}` + `tools/{server/{server-context.cpp,server-stream.cpp},mtmd/models/qwen3vl.cpp}` + `tests/**` + `conversion/**` + `gguf-py/**` + `tools/ui/**` | **Additive/tuning-only, no public-API surface (48 files, ~4.8k insertions incl. the auto-followed `tools/ui` WebUI + python conversion scripts, 16 commits).** All **eight** priority-8 headers (`common/{common.h,chat.h,speculative.h,arg.h,download.h}`, `tools/mtmd/mtmd.h`, `include/{llama.h,llama-cpp.h}`) are **byte-identical**. Two patch-target files changed but only **outside** the patched regions: `common/arg.cpp` adds speculative **sidecar** draft-repo resolution (MTP/DFlash/Eagle3 discovery in `common_models_handler_apply` — a `LLAMA_EXAMPLE_DOWNLOAD`-only region, well away from patch `0001`'s `common_params_parse`/`common_params_parse_main` block, and `common/arg.h` is unchanged); `tools/server/server-context.cpp` adds a null-`ctx_tgt` guard in `load_model` (return 400/false instead of a 500 crash) — a different sub-region than patch `0002`'s `load_progress_callback` guard and patch `0003`'s slot-prompt-similarity getters. `tools/server/server-stream.cpp` (compiled into `jllama`, not a patch target) gains 7 additive lines in `server_res_spipe::on_complete()`. The bulk is ggml backend work (CUDA GET_ROWS quantized-type support + topk-MoE, Vulkan queue/mutex refactor, WebGPU CONV_2D_DW, Hexagon, OpenVINO, kleidiai), new **Laguna** (poolside) model arch + Qwen3-VL mtmd + chat auto-parser tuning inside the upstream-compiled `llama`/`llama-common` libs, and python `conversion/`/`gguf-py/` tooling (not built here). No OuteTTS generator anchor touched (`tools/tts/tts.cpp` unchanged). All **six** patches (`0001`–`0003`, `0006`–`0008`) apply unchanged. b10092 is the topmost release at bump time. | +| b10076–b10092 | upstream verification (sandbox) | All **six** patches (`0001`–`0003`, `0006`–`0008`) re-verified against a clean b10092 checkout (ggml/llama.cpp commit `3ce7da2c8`): applied in filename order via `git apply`, all clean — the `b10076...b10092` diff's two patch-target files (`common/arg.cpp`, `tools/server/server-context.cpp`) changed only outside the patched regions (`0001`'s `common_params_parse*` block and `0002`/`0003`'s `load_model`/getter regions untouched), `0006`/`0007`'s `tools/server/server.cpp` and `0008`'s `tools/server/server-models.cpp` unchanged (`0006`'s standalone `git apply --check` fails only because its context is the post-`0001` tree, as documented). The OuteTTS generator ran clean against `tools/tts/tts.cpp @ b10092` (all anchors held, generated TU written). Per-platform build + `ctest` confirmation by the CI pipeline. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index ea8248c4a..6c7cda6e1 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b10076 + GIT_TAG b10092 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= @@ -196,7 +196,7 @@ execute_process( COMMAND ${CMAKE_COMMAND} -DTTS_SRC=${llama.cpp_SOURCE_DIR}/tools/tts/tts.cpp -DOUT_CPP=${JLLAMA_TTS_GEN_CPP} - -DLLAMA_TAG=b10076 + -DLLAMA_TAG=b10092 -P ${CMAKE_CURRENT_SOURCE_DIR}/cmake/generate-tts-upstream.cmake RESULT_VARIABLE JLLAMA_TTS_GEN_RESULT ) diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index 26d5b5801..32970b2f5 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b10076"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b10092"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b10076-"} — call + * plus the resolved upstream commit, e.g. {@code "b10092-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b10076"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b10092"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b10076"; + public static final String LLAMA_CPP_VERSION = "b10092"; // Constants holder — not instantiable. private LlamaCppVersion() {}