OSII builds a grounded, inspectable intelligence layer over a user's own files. People can browse, search, and chat with their corpus today; AI agents can use the same scopes, artifacts, provenance, and retrieval APIs for detailed future workflows.
The project also supports domain extension without a core fork. A subject matter expert can package an extractor, synthesizer, embedder, or enricher as a small container. For example, an experimental team can process thousands of run folders into a standard table artifact that is immediately visible in the dashboard and available to agents.
Read the documentation · Browse the documentation on GitHub · Follow the Python walkthrough
Choose the path that matches your goal:
- Use a corporate pilot release: follow Corporate pilot images and Quay releases. It runs the supported bundle from approved images.
- Develop or evaluate OSII from source: follow the steps below. This path runs editable code and reloads Python/dashboard changes.
- Build a processor: start with Extend OSII; a processor is an optional compute service, not a fork of core storage.
OSII uses Podman for packaged deployment and system-level development dependencies.
By default, OSII looks in this folder inside the downloaded project:
osii-data/
└── source/ Put the files and folders you want OSII to process here
On macOS or Linux, create it and initialize your settings with:
cd /path/to/osii
mkdir -p osii-data/source
cp .env.example .envOn Windows PowerShell:
cd C:\path\to\osii
New-Item -ItemType Directory -Force osii-data\source
Copy-Item .env.example .envYou can now drag files into the newly created osii-data/source folder using
Finder or File Explorer. Files placed in the repository root are intentionally
not shown; this avoids treating OSII's own code and configuration as your
corpus.
OSII mounts source read-only: it can read your originals but cannot modify or
delete them. The osii-data folder is ignored by Git, so your documents will
not accidentally be included in a commit.
You can instead use an existing folder anywhere on your computer. Open .env
in a text editor and set its absolute path:
OSII_SOURCE_DIR=/Users/your-name/Documents/my-filesOn Windows, use a forward-slash path such as:
OSII_SOURCE_DIR=C:/Users/your-name/Documents/my-filesA mounted shared or network drive works the same way; point
OSII_SOURCE_DIR at its mounted path (for example S:/project-files) and make
sure the account running OSII can read it. Intake intentionally browses only
inside this configured root rather than exposing the whole host filesystem.
Files added with the dashboard's Upload files button are stored separately.
Canonical extracted text and provenance stay as inspectable .osii files;
queue state and the rebuildable SQLite catalog stay under .osii/state/.
On macOS or Linux, start the editable stack:
cd /path/to/osii
make devmake dev requires no container runtime or preexisting model download. It runs the API
(including grounded chat), worker, MCP server, dashboard, and four independent Processor API
services directly from source: Python text-layer PDF/Office extraction, cited
source-excerpt previews that do not use AI, 384-dimensional lexical token/word-pair
hashing, and deterministic document statistics/frequent-keyword enrichment.
Scanned PDFs still require optional Tesseract OCR. The launcher checks ports and dependencies,
reloads backend services when source changes, and keeps generated data under
osii-data/.
OSII does not install or start the separate Ollama application. If you use
Ollama, manage it separately. OSII then
tries that separately running Ollama service with all-minilm for
semantic embeddings and Meta llama3.2:1b for chat and synthesis. Open
Tools & services → AI models to check the connection, see installed models, and explicitly download
either approved starter model when missing. OSII bundles neither Ollama nor
model weights. Use ollama list and ollama pull <model> for additional model
choices available in your environment. BM25 and extractive chat remain the automatic model-free
fallbacks. Lexical hashing is available as an explicit no-model vector option;
it is not presented as semantic search.
The Tesseract executable is also a separate manual installation in the
bare-metal workflow. Confirm tesseract --version works, then start OSII's
OpenCV/Tesseract wrapper with make dev-ocr-host or
.\scripts\osii.ps1 dev-ocr-host. Normal make dev does not start OCR.
On Windows PowerShell, use the equivalent launcher:
cd C:\path\to\osii
.\scripts\osii.ps1 devThe first development startup may take longer while dependencies are checked.
Later starts reuse the ignored osii-env Python environment. When the terminal
output settles, open:
- OSII dashboard: http://localhost:5173
- Backend status: http://localhost:8511/health
- Extractor docs: http://localhost:8092/docs
- Synthesizer docs: http://localhost:8093/docs
- Embedder docs: http://localhost:8085/docs
- Enricher docs: http://localhost:8094/docs
- Model-provider bridge docs: http://localhost:8095/docs
- Chat health: http://localhost:8511/api/chat/health
- MCP server: http://localhost:8022/mcp
In the dashboard, select Intake in the first sidebar section. Add files tests required tools and shows the extractor selected for each matched file type. Process library adds embeddings, summaries, enrichments, or a better extraction to existing documents without repeating unrelated work. Activity keeps run history out of the setup forms. The entire shared volume is the default scope; file-type and glob rules narrow it rather than replacing it. Files are processed sequentially and appear under Files as each extraction completes.
Re-extraction is versioned. A better extractor can be saved beside the current result or made primary while preserving the previous version. See Extraction versions and downstream lineage.
With an Ollama synthesis model selected, open a document's Wiki tab or an
individual collection to generate a grounded LLM wiki as a standard enrichment
artifact. The model runs outside OSII; the portable Markdown and provenance are
saved inside .osii. See Generate an LLM wiki.
The document Enrichments tab and collection Derived artifacts section also include model-free examples for a top-20 noun/adjective n-gram keyword table and a grounded named-entity candidate list. See Example keyword and entity enrichments.
Search and Chat retain up to 20 recent searches or prompts in the current browser so those pages remain useful between visits. Only the prompt, scope, time, and search mode are saved—never results, answers, or citations—and each entry or the complete browser-local history can be deleted. Saved root and collection keyword snapshots also provide one-click searches and grounded question starters. The Home page's collapsed Library Insights section exposes root-level wikis, tables, entity lists, and standard knowledge graphs without adding another service or storage authority.
Use Labels & Tags on a document to store portable sensitivity awareness, handling notes, and plain-text tags in its canonical object sidecar. These markings are metadata, not access controls. A collection can be exported as a manifest/checksum-validated OSII package and merged from Collections on another system; duplicate file IDs retain local data and union their labels. Delete File Data always shows the affected collections, folders, aggregate products, source file, and indexes before requiring exact confirmation. See Sensitive data, transfer, and deletion.
Open Tools & services before the first intake. Its Start & status, AI
models, Processing methods, and Custom services submenus state what
make dev started, what must be installed separately, and what each method
actually does.
For a domain
processor running on the host, use a base URL such as
http://127.0.0.1:8091; a packaged API container should use the processor's
Compose service name or host.containers.internal for a host service. Run
Health, then Test, and return to Intake to select Retest tools.
Keep the terminal window open while using OSII. Stop host processes with
Ctrl+C. There are no containers to stop after make dev.
Use the deployment-style stack when testing images rather than editing code:
make build
make runOn Windows PowerShell:
.\scripts\osii.ps1 build
.\scripts\osii.ps1 runmake run never rebuilds images. Use make containers-dev or
.\scripts\osii.ps1 containers-dev when you intentionally want to rebuild and
run the integrated stack with optional MCP and OCR. The normal run command
starts the eight application/baseline containers and does not silently require
optional MCP or OCR images.
The normal release publishes three OSII image artifacts: core (shared by API, worker, and chat), dashboard, and baseline processors. The extractor, synthesizer, embedder, enricher, and model-provider bridge remain separate containers but select commands from the same compact baseline image. See Publish OSII images to Quay.
make dev/make dev-host/.\scripts\osii.ps1 dev: run the complete editable development stack without containers.make dev-core/.\scripts\osii.ps1 dev-core: run application services without processors for external-integration testing.make dev-ollama/.\scripts\osii.ps1 dev-ollama: explicit alias for the normal Ollama-first development profile.make dev-corporate/.\scripts\osii.ps1 dev-corporate: prefer the built-in Shirty HTTP adapter, then Ollama, then extractive fallbacks.make dev-extractor,dev-synthesizer,dev-embedder, ordev-enricher(and matching PowerShell commands): run one processor independently.make dev-model-bridge/.\scripts\osii.ps1 dev-model-bridge: run only the HTTP-only Ollama/Shirty/OpenAI-compatible provider adapter.make dev-ocr-host/.\scripts\osii.ps1 dev-ocr-host: run the optional OpenCV/Tesseract OCR service directly on the host, with its tuning UI athttp://localhost:8080/demo.make dev-tika/.\scripts\osii.ps1 dev-tika: start only Apache Tika in Podman. Run it in a second terminal besidemake devto keep all OSII code editable on the host.make dev-containers/.\scripts\osii.ps1 dev-containers: run application services from source while Tika and Tesseract run in Podman for deployment parity.make dev-containers-insecure: the same hybrid workflow with Podman registry TLS verification disabled for image pulls and builds. On Windows use.\scripts\osii.ps1 dev-containers -InsecureRegistries. This is an explicit trust decision; prefer verified registry certificates when available.make dev-services/.\scripts\osii.ps1 dev-services: start only the Podman OCR services used bydev-containers.make dev-examples/.\scripts\osii.ps1 dev-examples: run editable OSII plus the example table enricher.make dev-all/.\scripts\osii.ps1 dev-all: include optional agents, OCR, and all example services in containers. Ollama remains separately managed.make containers-dev/.\scripts\osii.ps1 containers-dev: rebuild and run the normal deployment-style container stack.make build-release/.\scripts\osii.ps1 build-release: build the three normal publishable images exactly once each.make push-release/.\scripts\osii.ps1 push-release: push those three explicitly tagged images after a non-local registry prefix is supplied.make logs/.\scripts\osii.ps1 logs: follow service logs.make down/.\scripts\osii.ps1 down: stop the stack without deleting your data volume.make doctor/.\scripts\osii.ps1 doctor: report generated environments, model caches,node_modules, OSII data, and container storage without deleting anything.make catalog-rebuildandmake catalog-verify(with matching PowerShell commands): manage the disposable.osii/state/catalog.sqlite3read index.- Export components for separate corporate repositories.
Docker is supported as an override when it is your local container runtime:
make COMPOSE='docker compose' dev-containers.\scripts\osii.ps1 dev-containers -Runtime DockerSee the documentation index for three short starting paths. The remaining pages are task-oriented reference material; nobody needs to read the documentation tree from beginning to end. Common next steps are:
- Python module walkthrough
- Extend OSII
- Processor API v1
- Standard artifact formats
- Architecture
- Local operation
- Guaranteed local processors
- Sensitive data, transfer, and deletion
The repository currently includes:
- a backend for creating local OSII databases from source collections
- a REST API for serving OSII content, search, and derived artifacts
- a frontend for browsing and inspecting the resulting data
- supporting services for OCR and chat/RAG workflows
- a versioned processor SDK and copyable extension examples
At a high level, the system works in three stages:
- ingest a source file collection into a local structured OSII database
- serve that database through a backend API
- browse, inspect, search, and analyze the collection through the frontend
Browsing, lexical retrieval, hashing-vector retrieval, extractive grounded chat, baseline synthesis, and baseline enrichment run without a model connection. Tika and Tesseract remain optional OCR/deployment services. Ollama, OpenAI-compatible services, Shirty, and domain processors enhance capabilities without becoming dependencies of the basic user experience.
Retrieval defaults to sentence-aligned 768-character chunks with roughly 128 characters of overlap. Intake exposes these settings, while BM25 and semantic retrieval share the same chunk manifest and grounded source offsets. See retrieval chunking and overlap.
The repository remains under active development. Processor API v1 is the compatibility boundary for new extensions.
The monorepo currently contains components such as:
- OSII backend
- frontend dashboard / data viewer
- OCR service integrations
- chat / RAG support
- MCP and related tooling
Use the backend CLI to process a source collection into .osii.
For a comprehensive but small step-by-step walkthrough, use the
Jupytext-ready Python demonstrations. Plain
Python files are canonical; manage_notebooks.py converts the complete set to
or from notebooks without making notebook JSON the normal review format.
Run the backend FastAPI service to expose the OSII store over REST.
Run the frontend to browse and inspect the collection.
This is an open development repository. Expect:
- ongoing refactors
- evolving APIs
- incomplete documentation in some areas
- experimental features