ollama

mirror of https://github.com/jmorganca/ollama synced 2025-10-06 00:32:49 +02:00

Author	SHA1	Message	Date
Daniel Hiltgen	517807cdf2	perf: build graph for next batch async to keep GPU busy (#11863 ) * perf: build graph for next batch in parallel to keep GPU busy This refactors the main run loop of the ollama runner to perform the main GPU intensive tasks (Compute+Floats) in a go routine so we can prepare the next batch in parallel to reduce the amount of time the GPU stalls waiting for the next batch of work. * tests: tune integration tests for ollama engine This tunes the integration tests to focus more on models supported by the new engine.	2025-08-29 14:20:28 -07:00
Daniel Hiltgen	ed4e139314	Integration test improvements (#9654 ) Add some new test coverage for various model architectures, and switch from orca-mini to the small llama model.	2025-04-16 14:25:55 -07:00
CYJiang	e7019c9455	fix(integration): move waitgroup Add(1) outside goroutine to avoid potential issue (#10070 ) Signed-off-by: googs1025 <googs1025@gmail.com>	2025-04-08 15:17:40 -07:00
Bruce MacDonald	9876c9faa4	chore(all): replace instances of interface with any (#10067 ) Both interface{} and any (which is just an alias for interface{} introduced in Go 1.18) represent the empty interface that all types satisfy.	2025-04-02 09:44:27 -07:00
Stefan Weil	abfdc4710f	all: fix typos in documentation, code, and comments (#7021 )	2024-12-10 12:58:06 -08:00
Daniel Hiltgen	f0a351810c	tests: fix max queue integration test (#7782 ) This had fallen out of sync with the envconfig behavior, where max queue default was not zero.	2024-11-22 08:05:45 -08:00
Michael Yang	7ed367419e	fix concurrency test	2024-08-05 16:36:16 -07:00
Michael Yang	0f1910129f	int	2024-07-22 11:30:07 -07:00
Daniel Hiltgen	7f2fbad736	Skip max queue test on remote This test needs to be able to adjust the queue size down from our default setting for a reliable test, so it needs to skip on remote test execution mode.	2024-05-16 16:24:18 -07:00
Daniel Hiltgen	45d61aaaa3	Add integration test to push max queue limits	2024-05-05 10:46:25 -07:00

10 Commits