[pull] master from ggerganov:master #158

pull · 2024-12-02T13:54:20Z

See Commits and Changes for more details.

Created by pull[bot] (v2.0.0-alpha.1)

Can you help keep this open source service alive? 💖 Please sponsor : )

…q4_0_4x4_q8_0() (#10567) Signed-off-by: Adrien Gallouët <[email protected]>

* readme : update the usage section with examples * readme : more examples

* ggml : automatic selection of best CPU backend * amx : minor opt * add GGML_AVX_VNNI to enable avx-vnni, fix checks

…chat template types (#10572) * Templates: `mistral-v1`, `mistral-v2`, `mistral-v3`, `mistral-v3-tekken` * Changed system message logic and added tests for all 4 * Invalid `system_message` instead of `content` fixed * Removed tab-indented lines * Added template code and test for `mistral-v7` * Added all tests. Fixed bug with `tmpl == "llama2"` test. * Replaced tabs with spaces. * Removed `'mistral-v2'` option as no (open) models ever used it * Removed all references to 'v2' template from comments * Update llama.cpp Fixed `trim_assistant_message` bug

* contrib : refresh * contrib : expand [no ci] * contrib : expand test-backend-ops instructions * contrib : add CODEOWNERS * prs : update template to not have checkbox [no ci]

* Switched to GGML_LOG * Fix missing semicolon

* add cmake rvv support * add timings * remove space * update readme * fix * fix code * remove empty line * add test --------- Co-authored-by: Xuan Son Nguyen <[email protected]>

* make : deprecate ggml-ci * ci : disable Makefile builds ggml-ci * docs : remove make references [no ci] * ci : disable swift build ggml-ci * docs : remove obsolete make references, scripts, examples ggml-ci * basic fix for compare-commits.sh * update build.md * more build.md updates * more build.md updates * more build.md updates * Update Makefile Co-authored-by: Diego Devesa <[email protected]> --------- Co-authored-by: slaren <[email protected]>

* llama : add enum for supported chat templates * use "built-in" instead of "supported" * arg: print list of built-in templates * fix test * update server README

* server : force F16 KV cache for the draft model ggml-ci * server : fix draft params ggml-ci * server : various params fixes ggml-ci

this doesn't work as expected

* metal : small-batch mat-mul kernels ggml-ci * metal : add rest of types ggml-ci * metal : final adjustments ggml-ci * metal : add comments ggml-ci

* readme : document --no-display-prompt * readme : update default prompt context size * readme : remove unnecessary indentation Indenting a line with four spaces makes Markdown treat that section as plain text. * readme : indent commands under bullets * readme : indent commands in lettered list

* server : (refactoring) reduce usage of json internally * move all response types to struct * wip [no ci] * many fixes * add virtual function * fix index * minor style fix * add std::move * refactor handle_completions_generic * add virtual functions * remove server.hpp * clarify server_sent_event RFC specs * apply review comments * fix model_alias and completion_probabilities * small clean up * remove virtual for to_json_oai_compat() * naming oai_compat --> oaicompat * fix unwanted recursive call * update docs

* metal : Extend how Llama.cpp locates metal resources (#10675) * It searches the resource file in the directory where the current binary is located as well. * Resolves symbolic links. Rationale: When we plug this dependency into a Bazel build and run it in the context of Bazel (e.g. testing): * the execution directory is often very different from where the files are located and no direct control over this (Bazel sandboxing), * the Bazel sandbox often use symbolic links to make files available. With this patch, we can have the resource file added to the target, can build and run tests in the context of Bazel. * Update ggml/src/ggml-metal/ggml-metal.m Co-authored-by: Georgi Gerganov <[email protected]> * Update ggml/src/ggml-metal/ggml-metal.m Co-authored-by: Georgi Gerganov <[email protected]> --------- Co-authored-by: Georgi Gerganov <[email protected]>

…ng (#10597) * Vulkan: Implement VK_KHR_cooperative_matrix support in the matrix matrix multiplication shader * Improve performance with better q4_k and q5_k dequant and store unrolling * Add Vulkan MUL_MAT and MUL_MAT_ID accumulator precision selection * Rework mulmat shader selection and compilation logic, avoid compiling shaders that won't get used by device * Vulkan: Implement accumulator switch for specific mul mat mat shaders * Vulkan: Unroll more loops for more mul mat mat performance * Vulkan: Add VK_AMD_shader_core_properties2 support to read Compute Unit count for split_k logic * Disable coopmat support on AMD proprietary driver * Remove redundant checks * Add environment variable GGML_VK_DISABLE_COOPMAT to disable VK_KHR_cooperative_matrix support * Fix rebase typo * Fix coopmat2 MUL_MAT_ID pipeline selection

ggml-ci

* rename ggml-cpu-aarch64.c to .cpp * reformat extra cpu backend. - clean Q4_0_N_M and IQ4_0_N_M - remove from "file" tensor type - allow only with dynamic repack - extract cpu extra bufts and convert to C++ - hbm - "aarch64" - more generic use of extra buffer - generalise extra_supports_op - new API for "cpu-accel": - amx - aarch64 * clang-format * Clean Q4_0_N_M ref Enable restrict on C++ * add op GGML_OP_MUL_MAT_ID for Q4_0_N_M with runtime repack * added/corrected control on tensor size for Q4 repacking. * Update ggml/src/ggml-cpu/ggml-cpu-aarch64.cpp Co-authored-by: Georgi Gerganov <[email protected]> * Update ggml/src/ggml-cpu/ggml-cpu-aarch64.cpp Co-authored-by: Georgi Gerganov <[email protected]> * add debug logs on repacks. --------- Co-authored-by: Georgi Gerganov <[email protected]>

* server : various fixes ggml-ci * server : show curent seed in slot_params ggml-ci * fix /slots endpoint * Update examples/server/server.cpp Co-authored-by: Georgi Gerganov <[email protected]> * server : reflect endpoint response changes in the readme ggml-ci --------- Co-authored-by: Xuan Son Nguyen <[email protected]> Co-authored-by: Xuan Son Nguyen <[email protected]>

ggml-ci

* server : (refactor) no more json in server_task input * add test for slots endpoint * add tests for /props and /slots * remove task inf_type * fix CI by adding safe_json_to_str * add "model_path" to /props * update readme

* add 128k yarn context for Qwen * added property for model tensors * removing useless line

…10713)

* llama : use cmake for swift build * swift : <> -> "" * ci : remove make * ci : disable ios build * Revert "swift : <> -> """ This reverts commit d39ffd9. * ci : try fix ios build * ci : cont * ci : cont --------- Co-authored-by: Georgi Gerganov <[email protected]>

…10723) * Vulkan: fix NaN in tanh.comp * Faster NaN-free tanh

* server : bring back into to final chunk in stream mode * clarify a bit * traling space

* server : fix format_infill * fix * rename * update test * use another model * update test * update test * test_invalid_input_extra_req

…0668) * Update cmakepreset.json to use clang with ninja by default * Update cmakepreset.json to add clang and ninja based configs * Updates to build.md file * Make updates to rename preset targets * Update with .cmake file * Remove additional whitespaces * Add .cmake file for x64-windows-llvm * Update docs/build.md * Update docs/build.md --------- Co-authored-by: Max Krasnyansky <[email protected]>

There are some bugs in the 1.3.296 SDK, so disable this. It isn't strictly necessary anyway. Add missing dependency on vulkan-shaders-gen, so shaders get recompiled when it changes. Fix coopmat support reporting when glslc doesn't support NV_coopmat2.

Co-authored-by: eugenio.segala <[email protected]>

* Renames NVIDIA GPU-architecture flags to avoid name clashes with WinAPI. (e.g. CC_PASCAL, GPU architecture or WinAPI pascal compiler flag?) * Reverts erroneous rename in SYCL-code. * Renames GGML_CUDA_MIN_CC_DP4A to GGML_CUDA_CC_DP4A. * Renames the rest of the compute capability macros for consistency.

This allows for setting the --no-context-shift value in llama-imatrix which is required for models like DeepSeek

* q5_k q4_k q3_k q2_k q6_k multi row example * revert as multi row isnt faster for k quants

Vulkan doesn't mandate a specific rounding mode, but the shader_float_controls feature allows rounding mode to be requested if the implementation supports it.

* feat: load all backends from a user-provided search path * fix: Windows search path * refactor: rename `ggml_backend_load_all_in_search_path` to `ggml_backend_load_all_from_path` * refactor: rename `search_path` to `dir_path` * fix: change `NULL` to `nullptr` Co-authored-by: Diego Devesa <[email protected]> * fix: change `NULL` to `nullptr` --------- Co-authored-by: Diego Devesa <[email protected]>

angt and others added 11 commits November 30, 2024 09:13

ggml-cpu: replace AArch64 NEON assembly with intrinsics in ggml_gemv_…

0c39f44

…q4_0_4x4_q8_0() (#10567) Signed-off-by: Adrien Gallouët <[email protected]>

build: update Makefile comments for C++ version change (#10598)

43957ef

readme : update the usage section with examples (#10596)

6acce39

* readme : update the usage section with examples * readme : more examples

server : bind to any port when specified (#10590)

86dc11c

ggml : automatic selection of best CPU backend (#10606)

3420909

* ggml : automatic selection of best CPU backend * amx : minor opt * add GGML_AVX_VNNI to enable avx-vnni, fix checks

ci: add error handling for Python venv creation in run.sh (#10608)

5c7a5aa

grammars : add English-only grammar (#10612)

5e1ed95

contrib : refresh (#10593)

4cb003d

* contrib : refresh * contrib : expand [no ci] * contrib : expand test-backend-ops instructions * contrib : add CODEOWNERS * prs : update template to not have checkbox [no ci]

SYCL: Fix and switch to GGML_LOG system instead of fprintf (#10579)

991f8aa

* Switched to GGML_LOG * Fix missing semicolon

server: Add "tokens per second" information in the backend (#10548)

64ed209

* add cmake rvv support * add timings * remove space * update readme * fix * fix code * remove empty line * add test --------- Co-authored-by: Xuan Son Nguyen <[email protected]>

pull bot added the ⤵️ pull label Dec 2, 2024

github-actions bot added examples devops python server ggml SYCL testing build script labels Dec 2, 2024

github-actions bot added the documentation Improvements or additions to documentation label Dec 2, 2024

ngxson and others added 5 commits December 2, 2024 22:10

llama : add enum for built-in chat templates (#10623)

642330a

* llama : add enum for supported chat templates * use "built-in" instead of "supported" * arg: print list of built-in templates * fix test * update server README

server : fix default draft model parameters (#10586)

70b98fa

* server : force F16 KV cache for the draft model ggml-ci * server : fix draft params ggml-ci * server : various params fixes ggml-ci

github : minify link [no ci]

844e2e1

github : minify link [no ci] (revert)

515d4e5

this doesn't work as expected

metal : small-batch mat-mul kernels (#10581)

0115df2

* metal : small-batch mat-mul kernels ggml-ci * metal : add rest of types ggml-ci * metal : final adjustments ggml-ci * metal : add comments ggml-ci

github-actions bot added the Apple Metal label Dec 3, 2024

pminev and others added 29 commits December 5, 2024 22:36

fix(server) : not show alert when DONE is received (#10674)

7736837

common : bring back --no-warmup to server (#10686)

f162d45

convert : add custom attention mapping

c5ede38

convert : add support for Roberta embeddings (#10695)

784a14a

server : fix free of spec context and batch (#10651)

c2a16c0

ggml-ci

ggml : disable iq4_nl interleave size 8 (#10709)

d9c3ba2

ggml-ci

server : (refactor) no more json in server_task input (#10691)

3573fa8

* server : (refactor) no more json in server_task input * add test for slots endpoint * add tests for /props and /slots * remove task inf_type * fix CI by adding safe_json_to_str * add "model_path" to /props * update readme

llama : add 128k yarn context for Qwen (#10698)

62e84d9

* add 128k yarn context for Qwen * added property for model tensors * removing useless line

vulkan: compile a test shader in cmake to check for coopmat2 support (#…

ecc93d0

…10713)

Vulkan: fix NaN in tanh.comp with AMD proprietary driver on Windows (#…

06d7014

…10723) * Vulkan: fix NaN in tanh.comp * Faster NaN-free tanh

server : bring back info of final chunk in stream mode (#10722)

e52522b

* server : bring back into to final chunk in stream mode * clarify a bit * traling space

server : fix format_infill (#10724)

ce8784b

* server : fix format_infill * fix * rename * update test * use another model * update test * update test * test_invalid_input_extra_req

cmake : simplify msvc charsets (#10672)

1a05004

vulkan: fix compile warnings (#10731)

3d98b4c

CUDA: fix shared memory access condition for mmv (#10740)

26a8406

server : add flag to disable the web-ui (#10762) (#10751)

a86ad84

Co-authored-by: eugenio.segala <[email protected]>

imatrix : Add imatrix to --no-context-shift (#10766)

ae4b922

This allows for setting the --no-context-shift value in llama-imatrix which is required for models like DeepSeek

vulkan: dynamic subgroup size for the remaining k quants (#10745)

dafae66

* q5_k q4_k q3_k q2_k q6_k multi row example * revert as multi row isnt faster for k quants

vulkan: request round-to-even for fp16 in im2col/rope_head (#10767)

b685daf

Vulkan doesn't mandate a specific rounding mode, but the shader_float_controls feature allows rounding mode to be requested if the implementation supports it.

teleprint-me closed this Dec 11, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[pull] master from ggerganov:master #158

[pull] master from ggerganov:master #158

pull bot commented Dec 2, 2024 •

edited

Loading

[pull] master from ggerganov:master #158

[pull] master from ggerganov:master #158

Conversation

pull bot commented Dec 2, 2024 • edited Loading

pull bot commented Dec 2, 2024 •

edited

Loading