- Python 56.1%
- TypeScript 32.3%
- CSS 9.4%
- Dockerfile 1.6%
- HTML 0.4%
- Other 0.2%
|
|
||
|---|---|---|
| app | ||
| config/searxng | ||
| docs | ||
| .env.example | ||
| .gitignore | ||
| docker-compose.amd.yml | ||
| docker-compose.yml | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
X-Local-Search
A self-hosted, privacy-first search engine in the spirit of Brave Search and Perplexity — but everything runs on your machine. It combines a SearXNG metasearch backend with a local llama.cpp language model to give you:
- Research — the default search mode: an agentic, multi-step research pipeline. The model plans sub-queries, searches each (optionally scoped to a focus mode — general/news/science/social), ranks candidates by blended keyword + semantic relevance, reads the actual source pages, and streams back a cited markdown research report (live, token by token) — with a live sources preview that fills in as pages are read. Where sources genuinely disagree, the report says so explicitly instead of quietly picking one. An optional deep mode runs multiple rounds of search-and-read, reflecting after each on what's still missing before deciding whether to search again — minutes, not seconds, like real iterative deep-research tools — then re-checks its own citations against the sources, flagging anything unsupported.
- Ephemeral by default — no accounts, no sign-in. History is kept in SQLite only for the current session and wiped when you close the tab; you can opt into persisting it across restarts from Settings.
- Zero model setup — two fixed, fully open-weight models (a chat model and a small embedding model for reranking) are downloaded the first time you open the app, then loaded automatically forever after. No catalog, no picker, no decisions.
- GPU support — AMD (ROCm), or plain CPU.
The UI uses the Nord palette (#2e3440 base, #bf616a accent).
Architecture
Each variant is a single, self-contained compose file (CPU, AMD), with two services:
flowchart TD
U[Browser UI] --> X[xls service]
X -->|JSON API| S[SearXNG]
X -->|OpenAI API localhost| L[llama-server: chat model]
X -->|OpenAI API localhost| E[llama-server: embedding model]
X --> DB[(SQLite history + settings)]
X --> M[/GGUF models volume/]
S --> W[The web]
xls— the application. It is built FROM the officialllama.cppimage, so thellama-serverbinary and the correct GPU runtime are baked in. A small FastAPI backend orchestrates search, runs the research pipeline, manages two embeddedllama-serverprocesses (a chat model and a small CPU-only embedding model, each independently start/stop/idle-managed), fetches both models on first run, and serves the frontend.searxng— the official SearXNG image, configured with the JSON API enabled so the backend can query it programmatically.
How search works
Research (multi-step, streamed over SSE) is the default search mode, with an optional deep variant (see below):
- Plan — the model decomposes your topic into several focused sub-queries.
- Search — each sub-query is run through SearXNG, optionally scoped to a focus mode (general/news/science/social); candidate sources are collected and de-duplicated.
- Rank — candidates are scored by blended lexical + semantic relevance against the original topic (keyword overlap plus embedding cosine similarity from the small local embedding model — degrades gracefully to keyword-only if that model isn't ready), capped per-domain so one site can't crowd out the rest, and read in that order.
- Read — the top-ranked sources are fetched and their main content
extracted (via
trafilatura); pages that turn out to be off-topic once read are dropped in favor of the next-best candidate. - Synthesize — the model writes a structured report (summary first, then
sectioned detail, then key takeaways) citing every claim
[n]. Where sources genuinely disagree, it wraps the comparison in a:::disagreeblock naming which source says what, instead of silently picking one or inventing a synthesized middle ground. - Trim references — once the report is done, only the sources it actually cited become the References list; anything read but not cited is kept as a collapsed "Also reviewed" list instead of being discarded, so the primary links shown are the ones actually relevant to the answer.
- Related questions — a short follow-up call suggests a few related questions to keep exploring.
- Verify (deep mode only) — the model re-checks its own cited claims against the sources and flags anything unsupported or overstated as a "Confidence check" instead of silently leaving it in the report.
Deep mode isn't just "fast mode with bigger numbers" — it's an iterative
loop, matching how real deep-research tools actually work. After each
search-and-read round it reflects on what's been found and judges whether
real gaps remain; if so, it proposes new queries and runs another round, for
up to RESEARCH_DEEP_MAX_ROUNDS rounds (default 3) or until it judges the
research thorough enough. This takes minutes and reads well beyond a handful
of pages, on purpose — depth is the point.
Quick start
CPU
cp .env.example .env
docker compose up -d --build
AMD (ROCm)
Requires ROCm drivers and access to /dev/kfd + /dev/dri.
docker compose -f docker-compose.amd.yml up -d --build
Some cards need a gfx override — set HSA_OVERRIDE_GFX_VERSION in .env
(e.g. 10.3.0 for RX 6000 series).
First run
There is no wizard and no account. The first time you open the app it shows a single progress bar while it downloads the language model, then drops you straight into search. Every later visit opens instantly.
The models
Two models ship with the app — no picker, no catalog:
| Chat model | Embedding model | |
|---|---|---|
| Model | Qwen3 8B Instruct (Q4_K_M GGUF) |
bge-small-en-v1.5 (F16 GGUF) |
| Source | bartowski/Qwen_Qwen3-8B-GGUF |
CompendiumLabs/bge-small-en-v1.5-gguf |
| Download | ~5.0 GB, once | ~65 MB, once |
| Runs on | 6 GB VRAM fully offloaded (or CPU) | Always CPU — never competes with the chat model for VRAM |
| License | Apache 2.0 — open weights, no usage restrictions | MIT |
| Used for | Planning, synthesis, related questions, verification | Semantic reranking of search candidates only |
Both are Apache 2.0 / MIT, open weights, no usage restrictions.
Downloads once, stays on your machine.
Session lifecycle
History is ephemeral by default — no accounts, single-user:
- Searches are written to SQLite so you can re-open them from the sidebar during the session.
- Closing the tab (or reloading it) fires
POST /api/session/end(vianavigator.sendBeacon), which wipes the history unless persistence is on — see below. - History is also cleared on service start, so a crashed browser can't leave searches behind — again, unless persistence is on.
- The chat model unloads only after
IDLE_UNLOAD_SECONDSof real inactivity, not on every reload (a plain page refresh fires the same browser event as closing the tab, and reloading an already-downloaded model into memory is slow — so unload is tied to the idle timer, not page lifecycle). It reloads automatically on your next search. The embedding model is small/CPU-only and stays loaded.
Opt-in persistence: flip "Persist history" in Settings to keep history across restarts instead of wiping it. It's off by default. Clear-all/per-entry delete work the same either way — since there are no accounts, this is "this machine's data," not multi-user data.
Configuration
All options are set in .env (see .env.example). It's deliberately short —
most people only touch the AMD section if that applies to them:
| Variable | Default | Description |
|---|---|---|
APP_BIND / APP_PORT |
127.0.0.1 / 8080 |
Where the UI is exposed. Use 0.0.0.0 for LAN. |
DATA_DIR |
./data |
Host path for models + history DB. |
VIDEO_GID / RENDER_GID |
– | AMD only: host GIDs so the container can see the GPU. |
HSA_OVERRIDE_GFX_VERSION |
– | AMD only: gfx target override for cards ROCm doesn't recognize by default. |
LLAMA_CTX |
16384 |
Chat model context window. Lower it if you're low on VRAM/RAM. |
LLAMA_NGL |
0 (CPU file) / 99 (AMD file) |
GPU layers to offload; set 0 on the AMD file to force CPU for troubleshooting. |
RESEARCH_MAX_SOURCES |
0 |
Max sources read per search. 0 = no cap. |
IDLE_UNLOAD_SECONDS |
300 |
Unload the chat model after this idle time to free GPU/RAM (0 = never). Reloads on the next search. |
HF_TOKEN |
– | Optional Hugging Face token for gated/faster downloads. |
Research-pipeline tuning (sub-query count, ranking weights, etc.) has sane
defaults baked into app/backend/config.py and isn't surfaced as .env
options — if you want to change one, add it as an environment variable in
the compose file's environment: block (see the variable names in
config.py for what's available).
API
| Method | Endpoint | Purpose |
|---|---|---|
GET |
/api/health |
Liveness + model status. |
POST |
/api/deep-search |
Research run ({query, history, focus, mode, conversation_id}) — SSE event stream. |
GET |
/api/history |
List stored searches (this session, or persisted history if enabled). |
GET |
/api/history/{id} |
Full record for replay. |
DELETE |
/api/history/{id} |
Delete one entry. |
DELETE |
/api/history |
Clear the history. |
GET |
/api/settings |
Current settings (currently just persist_history). |
PUT |
/api/settings |
Update settings ({persist_history}). |
POST |
/api/session/end |
Tab closed or reloaded: wipe history unless persistence is on. |
GET |
/api/models/status |
Model state for both models (loaded / downloaded / GPU offload). |
GET |
/api/models/log |
Tail of the chat model's llama-server log. |
POST |
/api/models/bootstrap |
Download (if needed) + load both models — SSE progress. |
POST |
/api/models/stop |
Force-stop / unload the running chat model. |
Troubleshooting model loading
If the model downloads but fails to load, check GET /api/models/log (the
chat model's llama-server log tail) for the exact cause. Common ones:
unknown model architecture— the llama.cpp build in your image is older than the model. Rebuild pulling the latest base image:docker compose -f docker-compose.amd.yml build --pull, thenup -d.ROCm error/out of memory— the model is too big for your VRAM, or GPU init failed. The app now auto-retries on CPU and shows a ⚠ note; to force CPU setLLAMA_NGL=0in.env. For ROCm gfx mismatches, setHSA_OVERRIDE_GFX_VERSION(e.g.10.3.0,11.0.0).- Not enough VRAM — lower
LLAMA_CTX(e.g.4096) or setLLAMA_NGLto a partial layer count instead of99to split between GPU and CPU.
Notes & credits
- Inspired by odysseus (SearXNG-JSON + GPU compose patterns) and Brave Search / Perplexity's UX.
- SearXNG is pinned to a known-good tag; bump it deliberately after verifying a newer tag boots cleanly.
- This is a self-hosted tool with a powerful local model runtime. Keep it bound
to
127.0.0.1unless you know what you're exposing.