Context Extraction
OpenViking uses a three-layer async architecture for document parsing and context extraction.
Overview
Input File → Parser → TreeBuilder → SemanticQueue → Vector Index
↓ ↓ ↓
Parse & Move Files L0/L1 Generation
Convert Queue Semantic (LLM Async)
(No LLM)Design Principle: Parsing and semantics are separated. Parser doesn't call LLM; semantic generation is async.
Parser
Parser handles document format conversion and structuring, creating file structure in temp directory.
Supported Formats
| Format | Parser | Extensions | Status |
|---|---|---|---|
| Markdown | MarkdownParser | .md, .markdown | Supported |
| Plain text | TextParser | .txt | Supported |
| PDFParser | Supported | ||
| HTML | HTMLParser | .html, .htm | Supported |
| Code | CodeRepositoryParser | .py, .js, .go, etc. | Respects .gitignore and ignores common non-code directories |
| Image | ImageParser | .png, .jpg, etc. | |
| Video | VideoParser | .mp4, .avi, etc. | |
| Audio | AudioParser | .mp3, .wav, etc. |
Core Flow (Document Example)
# 1. Parse file
parse_result = registry.parse("/path/to/doc.md")
# 2. Returns temp directory URI
parse_result.temp_dir_path # viking://temp/abc123Smart Splitting
If document_tokens <= 1024:
→ Save as single file
Else:
→ Split by headers
→ Section < 512 tokens → Merge
→ Section > 1024 tokens → Create subdirectoryReturn Result
ParseResult(
temp_dir_path: str, # Temp directory URI
source_format: str, # pdf/markdown/html
parser_name: str, # Parser name
parse_time: float, # Duration (seconds)
meta: Dict, # Metadata
)TreeBuilder
TreeBuilder moves temp directory to AGFS and queues semantic processing.
Core Flow
building_tree = tree_builder.finalize_from_temp(
temp_dir_path="viking://temp/abc123",
scope="resources", # resources/user
)5-Phase Processing
- Find document root: Ensure exactly 1 subdirectory in temp
- Determine target URI: Map base URI by scope
- Recursively move directory tree: Copy all files to AGFS
- Clean up temp directory: Delete temp files
- Queue semantic generation: Submit SemanticMsg to queue
URI Mapping
| scope | Base URI |
|---|---|
| resources | viking://resources |
| user | viking://user |
SemanticQueue
SemanticQueue handles async L0/L1 generation and vectorization.
Message Structure
SemanticMsg(
id: str, # UUID
uri: str, # Directory URI
context_type: str, # resource/memory/skill
status: str, # pending/processing/completed
)Processing Flow (Bottom-up)
Leaf directories → Parent directories → RootSingle Directory Processing Steps
- Concurrent file summary generation: Limited to 10 concurrent
- Collect child directory abstracts: Read generated .abstract.md
- Generate .overview.md: LLM generates L1 overview
- Extract .abstract.md: Extract L0 from overview
- Write files: Save to AGFS
- Vectorize: Create Context and queue to EmbeddingQueue
Processing Limits
| Parameter | Default | Description |
|---|---|---|
max_concurrent_llm | 10 | Concurrent LLM calls |
max_images_per_call | 10 | Max images per VLM call |
max_sections_per_call | 20 | Max sections per VLM call |
Code Skeleton Extraction
For code files, OpenViking uses a fixed skeleton extraction route. This route is built into the code summary pipeline and is not selected or tuned by per-language parser settings.
What Skeleton Extraction Includes
The skeleton can include imports, classes, methods, functions, and other language-level symbols. Exact output depends on the maintained query or generic parser result for that language, but the route itself is fixed.
Extraction Route
Code skeleton extraction follows this fixed order:
- Use a maintained
tags.scmquery when one exists for the language. - If no corresponding
tags.scmexists, usetree-sitter-language-pack.process(). - Invoke
semantic.code_summaryonly as fallback when the extraction route produces no useful skeleton.
This routing applies to short and long code files alike.
Three Context Types Extraction
Flow Comparison
| Phase | Resource | Memory | Skill |
|---|---|---|---|
| Parser | Common flow | Common flow | Common flow |
| Base URI | viking://resources | viking://user/memories | viking://user/skills |
| TreeBuilder scope | resources | user | user |
| SemanticMsg type | resource | memory | skill |
Resource Extraction
# Add resource
await client.add_resource(
"/path/to/doc.pdf",
reason="API documentation"
)
# Flow: Parser → TreeBuilder(scope=resources) → SemanticQueueSkill Extraction
# Add skill
await client.add_skill({
"name": "search-web",
"content": "# search-web\\n..."
})
# Flow: Direct write to viking://user/skills/{name}/ → SemanticQueueMemory Extraction
# Memory auto-extracted from session
await session.commit()
# Flow: SessionCompressorV2 → ExtractLoop → MemoryUpdater → SemanticQueueRelated Documents
- Architecture Overview - System architecture
- Context Layers - L0/L1/L2 model
- Storage Architecture - AGFS and vector index
- Session Management - Memory extraction details
