Open source · Apache 2.0 Local-first · CPU-only No API key No billing Rust + Python

ThotBook-AI

The first fully scientist-made LLM toolset: an open-source, local, CPU-only pipeline that turns your own corpus — sessions, notebooks, paper indexes — into a private, OpenAI-compatible assistant. Reliable and FREE: free by principle, free by design, free as a common good.

Source code Documentation The research behind it

The pipeline — same tools, same maths, free

Ingest your knowledge corpus, tokenize it, then fine-tune and serve a private model — mirroring commercial "Copilot" products without the subscription, the data leaving your machine, or the vendor lock-in.

thotbook-ai pipeline: ingestion (M1 corpus ingest, M2 tokenize) feeds training (M3 LoRA, M4 RAG/HNSW) feeds serving (M5 OpenAI-compatible API, M7 llm-auto orchestration); the M6 citation crawler returns references to the corpus, and every stage reads and writes a versioned SQLite/Arrow/Delta lakehouse substrate.

Every stage writes versioned Apache Arrow → Delta lakehouse tables while redacting secrets at ingest time, and the whole pipeline runs on a laptop (32 GB RAM, no GPU). The companion scientific-computing engine optimiz-rs provides the underlying numerical primitives.

Local by design

Qwen3-8B served via Rust candle on CPU. No GPU, no cloud, no API key, no billing. The model never leaves your machine.

Private by principle

Every document is redacted at ingest time and a write-time leak audit refuses any row that still contains a secret. Your research stays yours.

Open by common good

Apache 2.0 weights, MIT / Apache 2.0 code, one scientist building research infrastructure for scientists — free on purpose, not as a trial.

OpenAI-compatible

Drop-in /v1/chat/completions. Ask from the CLI, from Python, or from opencode itself:

python -m thotbook_ai chat \
  "Give one line of Kähler geometry intuition."

Free for a reason.

Science advances when tools are shared. HFThot Research Lab publishes its working papers openly and builds the lab's toolset in the open so that every researcher — in any means-tested classroom, university, or start-up — has a free, trustworthy assistant rather than a vendor lease.

Built with and for the community

Contributions, bug reports and maths pedagogy feedback are welcome. The corpus pipeline, tokenizer and inference probes are public; the ncatlab-style inference benchmark ships with the repository.

research@hfthot-lab.eu GitHub