Files
obsidian_ollama/README.md
T
fegger 96f201bf3f Add dual-model support with separate chat and agent models
Split the single model setting into `chatModel` and `agentModel` to allow
using different LLMs for conversational modes (Ask, Research) versus
agentic modes (Edit, Organize, Workflow, auto-organizer). Defaults are
`deepseek-v4-flash` for chat and `glm-5.1` for agents.

Includes backward compatibility migration from legacy `model` field,
updated settings UI, per-mode tool filtering via new `agent-modes.ts`
configs, and vault search scoring improvements (exact phrase, recency,
filename bonuses).
2026-05-21 09:00:11 +02:00

17 KiB
Executable File
Raw Blame History

Obsidian Ollama Plugin

A plugin that integrates Ollama with Obsidian, allowing you to chat with local AI models, search your vault context, and use AI tools like creating files.

Features

  • Chat with Ollama models directly in Obsidian with streaming responses
  • Vault context search — the assistant can reference your notes via semantic (RAG) or keyword search
  • Agent Modes — selectable chat modes (Ask, Edit, Organize, Research, Workflow) that change available tools, system prompts, and preview behaviour
  • Tool integration — create, read, search, append, edit, rename, move, delete notes, and insert wiki-links
  • Structured Memory — persist conversation summaries, user preferences, and learned facts across sessions
  • Tool Telemetry — track which tools were called, which notes were searched, and LLM token usage
  • Semantic/RAG vault indexing — automatically index your vault into a vector database for intelligent retrieval
  • Semantic response cache — repeated or similar queries are answered instantly without hitting the model
  • Workflow Engine — execute multi-step AI workflows via /workflow commands
  • Auto-Organizer — AI-powered auto-tagging and auto-linking with dry-run preview and folder scoping
  • Obsidian MetadataCache integration — frontmatter, tags, links, and headings are read via Obsidian's built-in cache instead of raw regex parsing
  • Customisable model, URL, cache, and memory settings

Prerequisites

  1. Install Ollama: Follow the instructions at ollama.ai
  2. Start Ollama: ollama serve
  3. Pull a chat model: ollama pull llama3 (or any other model you prefer)

Optional — Vault Semantic Index (RAG)

The vault semantic index automatically indexes your Obsidian notes into a local ChromaDB vector database. When you ask a question, the plugin performs semantic search against your notes and includes the most relevant passages as context for the AI.

  1. Install ChromaDB:
    pip install chromadb
    
  2. Start ChromaDB:
    chroma run --host localhost --port 8000
    
  3. Pull an embedding model:
    ollama pull nomic-embed-text
    
  4. Enable the vault semantic index in the plugin settings and configure the ChromaDB URL.

Optional — Semantic Cache

The semantic cache stores responses in a local ChromaDB vector database. When you ask a question that is semantically similar to one already cached, the stored answer is returned immediately instead of calling the model.

  1. Install ChromaDB:
    pip install chromadb
    
  2. Start ChromaDB:
    chroma run --host localhost --port 8000
    
  3. Pull an embedding model (used to generate vectors for cache lookups):
    ollama pull nomic-embed-text
    
  4. Enable the cache in the plugin settings and configure the ChromaDB URL.

Installation

Use the included install script. It handles dependency installation, building, and copying the plugin into your vault:

# Clone or download this repository, then run:
./install.sh /path/to/your/obsidian/vault

The script will:

  • Install npm dependencies (excluding Ollama — you install that separately)
  • Compile the TypeScript plugin
  • Copy the built plugin into <vault>/.obsidian/plugins/ollama-plugin/

Manual install

If you prefer to install manually:

npm install
npm run build

Then copy the plugin into your vault:

mkdir -p /path/to/vault/.obsidian/plugins/ollama-plugin
cp manifest.json /path/to/vault/.obsidian/plugins/ollama-plugin/
cp main.js /path/to/vault/.obsidian/plugins/ollama-plugin/
cp styles.css /path/to/vault/.obsidian/plugins/ollama-plugin/
# Remove old dist/ from previous installs (no longer needed with bundling)
rm -rf /path/to/vault/.obsidian/plugins/ollama-plugin/dist

Note: The plugin is now bundled into a single main.js via esbuild. The obsidian npm package is a dev-only type stub — Obsidian provides its own API at runtime. The chromadb client library is also bundled into main.js, so no extra node_modules copy is needed for the semantic cache feature.

After installation

  1. Restart Obsidian (or reload: Ctrl+Shift+P → "Reload app without saving")
  2. Go to Settings → Community plugins → enable Ollama Plugin
  3. Configure the plugin at Settings → Ollama Settings

Configuration

Open Settings → Ollama Settings to configure the plugin.

Setting Default Description
Ollama URL http://localhost:11434 Base URL of your Ollama instance
Chat Model deepseek-v4-flash Model used for normal chat, Ask mode, and Research mode
Agent Model glm-5.1 Model used for Edit, Organize, Workflow, and auto-organizer tasks
Default Agent Mode Ask Default chat mode (Ask, Edit, Organize, Research, Workflow)
Vault Search Limit 5 Maximum number of vault entries to include in context
Max Context Length 8000 Maximum characters of vault content sent to the AI per message
Max Message History 50 Maximum number of messages kept in conversation history
Enable Vault Semantic Index Off Index vault notes into a vector DB for semantic/RAG search
Vault Index ChromaDB URL http://localhost:8000 URL of your ChromaDB instance for the vault index
Vault Index Embedding Model nomic-embed-text Ollama model used to generate vault embeddings
Vault Index Similarity Threshold 0.75 Minimum cosine similarity (01) for a vault search hit
Rebuild Vault Index Button to rebuild the entire vault semantic index
Clear Vault Index Button to delete all indexed vault notes
Enable Semantic Cache Off Cache responses for fast repeated queries
ChromaDB URL http://localhost:8000 URL of your running ChromaDB instance
Cache Embedding Model nomic-embed-text Ollama model used to generate cache embeddings
Cache Similarity Threshold 0.85 Minimum cosine similarity (01) for a cache hit
Clear Semantic Cache Button to wipe all cached responses
Enable Auto-Tagging Off Automatically suggest and apply tags to untagged notes
Max Tags Per Note 5 Maximum tags to generate per note
Normalize Tags On Normalize generated tags against existing vault vocabulary
Target Folder (Auto-Tag) Restrict auto-tagging to a specific folder
Enable Auto-Linking Off Add "Related Notes" sections based on semantic similarity
Max Links Per Note 3 Maximum related note links to insert
Target Folder (Auto-Link) Restrict auto-linking to a specific folder
Dry Run Mode (Auto-Link) Off Preview proposed links without applying them
Enable Structured Memory On Inject remembered context from past sessions into prompts
Max Conversation Summaries 10 Maximum past conversation summaries to retain
Max User Preferences 20 Maximum user preferences to retain
Max Learned Facts 50 Maximum learned facts to retain
Clear Structured Memory Button to delete all stored memory
Enable Tool Telemetry On Record tool calls, searches, and LLM token counts
Max Telemetry Entries 100 Maximum telemetry events to retain
Clear Tool Telemetry Button to delete all recorded telemetry

Usage

  1. Open the chat view via the command palette (Ctrl+P → "Open Ollama Chat") or the ribbon icon
  2. Type your message in the input box
  3. Press Enter or click Send to send your message
  4. Press Shift+Enter to insert a line break
  5. Click New Chat to start a fresh conversation

Agent Modes

The chat view includes a mode selector dropdown. Each mode changes the assistant's behaviour:

Mode Tools Available Preview Required Use Case
Ask Read, Search No Answer questions using vault context
Edit All tools Yes Create, modify, and manage notes
Organize Read, Search, Frontmatter, Rename, Move, Link Yes Tag, rename, move, and link notes
Research Read, Search No Deep vault search and synthesis
Workflow None (uses /workflow) No Execute multi-step AI workflows

When a mode requires preview (Edit, Organize), write operations like create_note or delete_note show a card with a before/after diff and Apply / Cancel buttons. Ask and Research modes execute write tools immediately without preview.

Workflows

Type /workflow followed by a description to trigger the workflow engine. The AI will generate a multi-step workflow plan, then execute it step-by-step. Example:

/workflow Find all notes tagged "meeting", summarise them, and create a "Meeting Summary" note

Auto-Organizer

Use the command palette to trigger:

  • Auto-Tag Untagged Notes — AI generates tags for notes missing tags
  • Auto-Link Related Notes — AI inserts "Related Notes" sections with wiki-links

Both features support:

  • Dry-run mode — preview proposed changes without modifying the vault
  • Target folder — restrict processing to a specific folder and its subfolders
  • Tag normalisation — match generated tags against existing vault vocabulary

Commands

Command Description
Open Ollama Chat Open the chat sidebar
Clear Semantic Cache Delete all cached responses
Clear Vault Index Delete all indexed vault notes
Rebuild Vault Index Rebuild the vault semantic index from scratch
Auto-Tag Untagged Notes Run the auto-tagger
Auto-Link Related Notes Run the auto-linker
Clear Structured Memory Delete all conversation summaries, preferences, and facts
Clear Tool Telemetry Delete all recorded telemetry events

Semantic Cache Behaviour

  • The cache is bypassed when tool calls are involved (e.g. file creation), since those requests have side effects.
  • Responses are stored against the last user message in the conversation. If a new query is sufficiently similar (above the configured threshold), the cached response is returned.
  • Re-asking the same question updates the existing cache entry rather than creating a duplicate.
  • Use the Clear Semantic Cache button in settings to remove all stored responses (for example after switching embedding models).

Vault Context

When you send a message, the plugin searches your vault for relevant notes and includes them as context. If the Vault Semantic Index is enabled, search is performed via semantic/RAG retrieval using vector embeddings. Otherwise, it falls back to a weighted keyword search:

  • Headings — 5x weight
  • Frontmatter title — 3x weight
  • Frontmatter tags — 2.5x weight
  • First paragraph — 1.5x weight
  • General content — 1x weight

The plugin also pulls in:

  • Explicit mentions — notes referenced via [[...]] wikilinks in the message
  • Open note — the currently active note
  • Selected text — text selected in the active editor
  • Backlinks / Outlinks — notes that link to / from the open note
  • Related notes — semantically similar notes (requires vault semantic index)

The plugin automatically watches your vault for changes (create, modify, delete, rename) and updates the semantic index in real time when enabled.

Frontmatter, tags, links, and headings are resolved using Obsidian's built-in metadataCache API for accuracy and performance.

Tools

The assistant has access to a suite of tools that interact with your vault. Available tools depend on the current Agent Mode:

Tool Description Mode
create_note / create_file Create a new markdown file Edit
read_vault_file Read the contents of a note All
search_vault_files Keyword-search vault files by path All
append_to_note Append text to the end of a note Edit
replace_note_section Replace content under a specific heading Edit
update_frontmatter Add, update, or remove frontmatter fields Edit, Organize
rename_note Rename a note file Edit, Organize
move_note Move a note to a different folder Edit, Organize
delete_note Delete a note Edit
insert_link Insert a [[wiki-link]] into a note Edit, Organize

Paths are validated for safety: no .obsidian/.git access, no path traversal (..), no absolute paths, and a 200-character limit.

Structured Memory

When Enable Structured Memory is on, the plugin remembers context across sessions by storing three kinds of data in Obsidian's plugin data JSON:

  • Conversation Summaries — After each assistant reply, a brief summary (topic + key points) is saved
  • User Preferences — Statements like "I prefer dark mode" or "My favourite colour is blue" are extracted and stored
  • Learned Facts — Simple facts mentioned in conversation (e.g., "Obsidian is a note-taking app") and vault folder paths are remembered

These are injected as a system message at the start of every LLM call, so the assistant "remembers" context from previous sessions. Limits and clear controls are available in settings.

Tool Telemetry

When Enable Tool Telemetry is on, the plugin records:

  • Tool calls — which tool, arguments, success/failure, result summary, and duration
  • LLM calls — model, estimated prompt/completion/total tokens, and duration
  • Vault searches — query, number of results, and matched note paths

Telemetry is stored locally in Obsidian's plugin data. The settings tab shows a Recent Activity summary of the last 10 events. Use Clear Tool Telemetry to wipe the history.

Note: Token counts are exact when Ollama provides prompt_eval_count and eval_count in its response; otherwise they are estimated from character count (÷4 approximation).

Supported Models

Any Ollama-supported model works. Popular choices:

Development

npm install
npm run build
npm test          # 520+ unit tests across 21 test suites
npm run lint      # ESLint check

The project uses TypeScript, Jest, and esbuild. Obsidian APIs are mocked in __mocks__/obsidian.ts for testing.

Troubleshooting

Symptom Likely cause Fix
Plugin doesn't appear in Obsidian Install script was not run or failed Run ./install.sh /path/to/vault and reload Obsidian
Cannot connect to Ollama Ollama is not running Run ollama serve
Model not found Model not pulled Run ollama pull <model>
"Invalid response format" error Ollama returned a non-JSON response (e.g., proxy error page) Check that Ollama is healthy at the configured URL
Semantic cache unavailable (notice shown) ChromaDB is not running, or the ChromaDB URL is wrong Start ChromaDB (chroma run) and verify the URL in settings
Cache always misses Similarity threshold is too high, or the embedding model was changed Lower the threshold or click Clear Semantic Cache and let the cache rebuild
Slow first response after enabling cache Embedding model not yet pulled Run ollama pull nomic-embed-text (or the model you configured)
Permission issues Vault write permissions Check that your Obsidian vault has proper write permissions
Structured memory not showing up Memory was just cleared or is empty Have a few conversations — summaries are generated after each assistant reply
Tool telemetry not recording Telemetry is disabled or max entries is 0 Enable Tool Telemetry in settings and set Max Telemetry Entries > 0

Security

  • File paths are validated to prevent access to .obsidian/ and .git/ directories
  • Path traversal attempts (..) are blocked
  • Absolute paths and Windows drive letters are rejected
  • Maximum path length is enforced (200 characters)

License

MIT License

Copyright (c) 2024 Flo Egger

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.