* server: don't walk Windows junctions in file_glob_search std::filesystem reports a junction as a plain directory, so the symlink guard misses it and a junction pointing back at an ancestor is walked until the path length gives out read the reparse tag and treat a symlink and a mount point as links, leaving any other reparse point walkable so cloud placeholders and dedup stubs still get searched look junk directory names up case insensitively on Windows, where NTFS makes Build the same directory as build test that a junk directory stays selectable while its contents stay out of search results * server: report a directory the walk could not read a directory that fails to open or to iterate was skipped in silence, so a caller got a listing that looked complete while a whole subtree was missing: a path over the platform limit, a volume going away, a name the filesystem rejects skip_permission_denied never reaches this path, so an error here is an incomplete answer rather than a deliberate omission, and it now sets the truncated flag * server: simplify the file_glob_search listing plumbing return a small result struct instead of two out params and a caller path that only fed an error string, taking list_entries from six parameters down to three scope the error code to the directory being read, act on the status code the entry lookups already returned, and treat an unreadable link state as a link so the walk never descends on a guess check the deadline when a directory is popped, not only per entry, so a tree of empty directories cannot outlive the budget read the path parameter once, and reject an invalid limit the way an invalid type is already rejected, instead of silently falling back normalize the resolved path, so a "." or ".." a caller typed reaches neither git nor the client, and return the generic path form with '/' separators on every platform, so the base sent to clients no longer needs a local fixup * ui: expire cached picker searches the cache grew for the lifetime of the component: entries went stale after the TTL but were never removed, so every distinct query typed in a session stayed in memory drop expired entries when a new result is stored * server: address review from @ngxson trim comments to one line each, and drop two that restate the code rename junk_lookup_name to get_effective_name, and move it and the link check to private static members next to junk_dir_names merge the Windows and Linux link checks into one is_link, so symlinks are checked everywhere and junctions only add to it on Windows * server: convert tool paths as UTF-8 on Windows a narrow path uses the active code page there, so a file name came back mangled and a path with an accent could not be opened at all convert explicitly at every crossing between a std::string, which always carries UTF-8 here, and fs::path read the home directory through the wide environment, since the narrow one returns the profile path in the active code page too the walker no longer normalizes separators by hand, since paths now come back in generic form * server: fold the platform branch inside console_output_to_utf8 match the shape of the other helpers, one definition with the #if inside, instead of two definitions wrapped in #if and #else inline the single caller helper and trim the comment
Server tests
Python based server tests scenario using pytest.
Tests target GitHub workflows job runners with 4 vCPU.
Note: If the host architecture inference speed is faster than GitHub runners one, parallel scenario may randomly fail.
To mitigate it, you can increase values in n_predict, kv_size.
Install dependencies
pip install -r requirements.txt
Run tests
- Build the server
cd ../../..
cmake -B build
cmake --build build --target llama-server
- Start the test:
./tests.sh
It's possible to override some scenario steps values with environment variables:
| variable | description |
|---|---|
PORT |
context.server_port to set the listening port of the server during scenario, default: 8080 |
LLAMA_SERVER_BIN_PATH |
to change the server binary path, default: ../../../build/bin/llama-server |
DEBUG |
to enable steps and server verbose mode --verbose |
N_GPU_LAYERS |
number of model layers to offload to VRAM -ngl --n-gpu-layers |
LLAMA_CACHE |
by default server tests re-download models to the tmp subfolder. Set this to your cache (e.g. $HOME/Library/Caches/llama.cpp on Mac or $HOME/.cache/llama.cpp on Unix) to avoid this |
To run slow tests (will download many models, make sure to set LLAMA_CACHE if needed):
SLOW_TESTS=1 ./tests.sh
To run with stdout/stderr display in real time (verbose output, but useful for debugging):
DEBUG=1 ./tests.sh -s -v -x
To run all the tests in a file:
./tests.sh unit/test_chat_completion.py -v -x
To run a single test:
./tests.sh unit/test_chat_completion.py::test_invalid_chat_completion_req
Hint: You can compile and run test in single command, useful for local development:
cmake --build build -j --target llama-server && ./tools/server/tests/tests.sh
To see all available arguments, please refer to pytest documentation
Debugging external llama-server
It can sometimes be useful to run the server in a debugger when invesigating test
failures. To do this, the environment variable DEBUG_EXTERNAL=1 can be set
which will cause the test to skip starting a llama-server itself. Instead, the
server can be started in a debugger.
Example using gdb:
$ gdb --args ../../../build/bin/llama-server \
--host 127.0.0.1 --port 8080 \
--temp 0.8 --seed 42 \
--hf-repo ggml-org/models --hf-file tinyllamas/stories260K.gguf \
--batch-size 32 --no-slots --alias tinyllama-2 --ctx-size 512 \
--parallel 2 --n-predict 64
And a break point can be set in before running:
(gdb) br server.cpp:4604
(gdb) r
main: server is listening on http://127.0.0.1:8080 - starting the main loop
srv update_slots: all slots are idle
And then the test in question can be run in another terminal:
(venv) $ env DEBUG_EXTERNAL=1 ./tests.sh unit/test_chat_completion.py -v -x
And this should trigger the breakpoint and allow inspection of the server state in the debugger terminal.