Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 9 additions & 11 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,15 +36,16 @@ ast-rag callers <name> --depth 2
| `sig` | `ast-rag sig "process(int, String)"` | Signature search |
| `evaluate` | `ast-rag evaluate --all` | Check quality |
| `index-folder` | `ast-rag index-folder ./ast_rag` | Index a folder |
| `workspace` | `ast-rag workspace . --apply` | Update from git diff |
| `workspace` | `ast-rag workspace . --apply` | Update from uncommitted changes |
| `stats` | `ast-rag stats` | See what is actually indexed |

### Python API

```python
from ast_rag.ast_rag_api import ASTRagAPI
from ast_rag.models import ProjectConfig
from ast_rag.graph_schema import create_driver
from ast_rag.embeddings import EmbeddingManager
from ast_rag.repositories import create_driver
from ast_rag.services import EmbeddingManager

# Initialize
cfg = ProjectConfig()
Expand Down Expand Up @@ -146,11 +147,8 @@ curl http://localhost:6333/collections
**Check index:**

```bash
# How many nodes in graph
cypher-shell "MATCH (n) RETURN count(n)"

# How many files indexed
grep "COMPLETE" /tmp/index_*.log | wc -l
# Node/edge counts, language distribution, indexed file count
ast-rag stats
```

---
Expand Down Expand Up @@ -191,10 +189,10 @@ ast-rag evaluate --all

```bash
# Check what's indexed
grep "COMPLETE" /tmp/index_*.log | wc -l
ast-rag stats

# Index remaining
./scripts/index-remaining.sh
# Index what is missing
ast-rag index-folder ./path/to/folder --no-schema

# Run again
ast-rag evaluate --all
Expand Down
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,7 @@ pytest tests/ -v

## 🎯 Areas Needing Contribution

- [ ] More language support (Go, Kotlin, Swift)
- [ ] More language support (C#, Kotlin, Swift)
- [ ] Performance optimizations
- [ ] Additional MCP tools
- [ ] Better error handling
Expand Down
81 changes: 64 additions & 17 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,20 @@ export AST_RAG_QDRANT_URL=http://myhost:6333
Or copy `ast_rag_config.example.json` to `ast_rag_config.json` and edit it
(it is gitignored).

### Optional extras

`pip install -e .` gives you the `ast-rag` CLI. The other two console scripts
need extras:

```bash
pip install -e ".[mcp]" # ast-rag-mcp (MCP server), ast-rag-watch (file watcher)
pip install -e ".[server]" # embedding_server.py, for the GPU host
pip install -e ".[dev]" # pytest, mypy, ruff
```

Without `[mcp]`, `ast-rag-mcp` and `ast-rag-watch` are on `PATH` but fail with
`ModuleNotFoundError: No module named 'mcp'` / `'watchdog'`.

## ⚙️ Configuration

Create `ast_rag_config.json` in project root:
Expand All @@ -75,12 +89,24 @@ Create `ast_rag_config.json` in project root:
"collection_name": "ast_rag_nodes"
},
"embedding": {
"model_name": "bge-m3",
"remote_url": "http://localhost:1113/v1/embeddings"
"model_name": "BAAI/bge-m3"
}
}
```

Embeddings are computed locally by default. To offload them to a separate
embedding server, add `remote_url` — and `dimension` with it, which is required
whenever `remote_url` is set (`docker compose up -d` does not start such a
server):

```json
"embedding": {
"model_name": "bge-m3",
"remote_url": "http://localhost:1113/v1/embeddings",
"dimension": 1024
}
```

## 🎯 Quick Start

```bash
Expand Down Expand Up @@ -116,30 +142,51 @@ ast-rag evaluate --all
## 🛠️ CLI Commands

```
ast-rag init <path> # Full indexing
ast-rag update <path> # Update from git diff
ast-rag query "<text>" # Semantic search
ast-rag goto <name> # Find definition
ast-rag callers <name> # Find callers
ast-rag refs <name> # Find references
ast-rag sig <pattern> # Signature search
ast-rag evaluate # Quality evaluation
ast-rag index-folder <path> # Index a folder
ast-rag workspace <path> # Show workspace changes
ast-rag sandbox <lang> <cmd> # Run in Docker sandbox
ast-rag init <path> # Full indexing
ast-rag index-folder <path> # Index a folder
ast-rag create-database <name> # Create a Neo4j database
ast-rag update <path> --from-commit <a> --to-commit <b>
# Update from git diff a..b
ast-rag workspace <path> # Show workspace changes
ast-rag query "<text>" # Semantic search
ast-rag goto <name> # Find definition
ast-rag callers <name> # Find callers
ast-rag refs <name> # Find references
ast-rag call-graph <name> # Callers/callees graph around a function
ast-rag symbol-impact <name> # Definition + refs + callers + callees
ast-rag sig <pattern> # Signature search
ast-rag blocks <function> # if/for/while/try/lambda/with blocks
ast-rag lambdas # List lambda/closure blocks
ast-rag summarize <name> # LLM summary of a function/method/class
ast-rag analyze-stacktrace [file] # Map a stack trace onto the graph
ast-rag stats # Statistics about the indexed codebase
ast-rag cache-stats # Parse cache statistics
ast-rag evaluate # Quality evaluation
ast-rag sandbox <lang> [workdir] --cmd "<command>"
# Run in Docker sandbox
```

`ast-rag <command> --help` lists the flags for each.

## 📊 Quality

Current metrics (Phase 2):
Phase 2 benchmark run, recorded 2026-02-26 and not re-measured since:

| Metric | Target | Actual |
| Metric | Target | Measured |
|--------|--------|--------|
| **Pass Rate** | >80% | **100%** ✅ |
| **F1 Score** | >0.85 | **0.98** ✅ |
| **Precision** | >0.85 | **0.98** ✅ |
| **Recall** | >0.85 | **0.97** ✅ |

The benchmarks in `benchmarks/queries/` are ground-truthed against AST-RAG's own
source, so they measure this repo rather than yours. To reproduce, index this
repo with Neo4j and Qdrant running, then from the repo root:

```bash
ast-rag evaluate --all
```

## 🏗️ Architecture

```
Expand Down Expand Up @@ -201,9 +248,9 @@ Planned features and improvements:

| Feature | Status | Description |
|---------|--------|-------------|
| **Code Summaries** | 🔜 Planned | Generate AI-powered summaries for functions/classes |
| **Code Summaries** | ✅ Done | `ast-rag summarize` — LLM summaries for functions/classes ([docs](docs/SUMMARIZATION.md)) |
| **Refactoring Hints** | 🔜 Planned | Detect code smells and suggest improvements |
| **More Languages** | 🔜 Planned | Go, C#, Kotlin with full AST support |
| **More Languages** | 🔜 Planned | C#, Kotlin with full AST support |
| **IDE Integration** | 🔄 In Progress | MCP, skills, CLI for OpenCode, Kilocode, Claude Code, Cursor |
| **Incremental Indexing** | ✅ Done | Git-based and filesystem watcher updates (improving for large codebases) |
| **Rust Rewrite** | 🔜 Planned | Full rewrite in Rust for performance, type safety, and easier integration |
Expand Down
30 changes: 19 additions & 11 deletions docs/QUICKSTART.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,12 +69,18 @@ cat > ast_rag_config.json <<EOF
},
"embedding": {
"model_name": "bge-m3",
"remote_url": "http://localhost:1113/v1/embeddings"
"remote_url": "http://localhost:1113/v1/embeddings",
"dimension": 1024
}
}
EOF
```

`remote_url` is what routes embeddings to the server from step 2 — drop it to
embed locally instead. When it is set, `dimension` must be set too, otherwise
collection creation fails with
`EmbeddingConfig.dimension must be set when using remote_url`.

### 2. Index project

```bash
Expand All @@ -96,7 +102,10 @@ ast-rag index-folder ./src
ast-rag evaluate --all
```

**Expected result:**
The benchmarks live in `benchmarks/queries/` and are ground-truthed against
AST-RAG's own source, so this only means anything with the AST-RAG repo indexed,
and it has to be run from the repo root. Output looks like:

```
📊 Benchmarks: 10
✅ Passed: 10
Expand Down Expand Up @@ -164,12 +173,11 @@ ast-rag workspace . --apply

```bash
# Update from git diff
ast-rag update . --from HEAD~1 --to HEAD

# Update current branch
ast-rag update . --current-branch
ast-rag update . --from-commit HEAD~1 --to-commit HEAD
```

Both `--from-commit` and `--to-commit` are required.

---

## 📊 Typical Scenarios
Expand Down Expand Up @@ -253,11 +261,11 @@ ast-rag init /path/to/codebase
### Low quality (<70%)

```bash
# Check how many indexed
grep "COMPLETE" /tmp/index_*.log | wc -l
# Check what is actually in the graph
ast-rag stats

# Index remaining
./scripts/index-remaining.sh
# Re-index the folders that are missing
ast-rag index-folder ./path/to/folder --no-schema

# Run evaluation again
ast-rag evaluate --all
Expand Down Expand Up @@ -301,7 +309,7 @@ ast-rag query "main initialization entry point" --limit 5
ast-rag goto <found_name> --snippet

# 3. Find what it calls (call graph)
ast-rag callers <entry_point> --depth 2
ast-rag call-graph <entry_point> --direction callees --depth 2

# 4. Find key dependencies
ast-rag query "database connection manager" --limit 5
Expand Down
4 changes: 2 additions & 2 deletions docs/STACK_TRACE_MAPPING.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,8 @@ ast-rag analyze-stacktrace error.log -v

```python
from ast_rag.stack_trace import StackTraceService
from ast_rag.graph_schema import create_driver
from ast_rag.embeddings import EmbeddingManager
from ast_rag.repositories import create_driver
from ast_rag.services import EmbeddingManager
from ast_rag.models import ProjectConfig

# Инициализация
Expand Down
10 changes: 5 additions & 5 deletions docs/SUMMARIZATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,13 +51,13 @@ ast-rag summarize MyFunction --max-callers 10 --max-callees 10

```python
from ast_rag.ast_rag_api import ASTRagAPI
from ast_rag.graph_schema import create_driver
from ast_rag.embeddings import EmbeddingManager
from ast_rag.repositories import create_driver
from ast_rag.services import EmbeddingManager
from ast_rag.summarizer import SummarizerService, NodeSummary
from ast_rag.models import ProjectConfig

# Load configuration
cfg = ProjectConfig.model_validate_json("ast_rag_config.json")
# Load configuration (model_validate_json takes JSON text, not a path)
cfg = ProjectConfig.model_validate_json(open("ast_rag_config.json").read())

# Initialize API
driver = create_driver(cfg.neo4j)
Expand Down Expand Up @@ -187,7 +187,7 @@ Processes HTTP requests by validating input, calling external services, and retu
|-----------|---------|-------------|
| `base_url` | `http://localhost:11434/v1` | OpenAI-compatible API endpoint |
| `model` | `qwen2.5-coder:14b` | Model name for summarization |
| `api_key` | `ollama` | API key (not needed for Ollama) |
| `api_key` | `None` | API key (not needed for Ollama) |
| `timeout` | `120` | Request timeout in seconds |

### Supported LLM Backends
Expand Down
18 changes: 9 additions & 9 deletions docs/agent-scenarios.md
Original file line number Diff line number Diff line change
Expand Up @@ -154,8 +154,8 @@ ast-rag evaluate --all
# Index specific folder
ast-rag index-folder ./src/modified_module --no-schema

# Or all remaining
./scripts/index-remaining.sh
# Or re-index everything
ast-rag init .
```

---
Expand All @@ -172,8 +172,8 @@ ast-rag sig "process(int, String)"
ast-rag sig "get*" --lang java
ast-rag sig "*Handler" --lang python

# With filter
ast-rag sig "build*" --lang rust --kind Function
# With a result cap
ast-rag sig "build*" --lang rust --limit 50
```

**Python API:**
Expand Down Expand Up @@ -203,11 +203,11 @@ ast-rag evaluate --query benchmarks/queries/def_001.json

**If quality is low:**
```bash
# Check how many indexed
grep "COMPLETE" /tmp/index_*.log | wc -l
# Check what is in the graph
ast-rag stats

# Index remaining
./scripts/index-remaining.sh
# Index what is missing
ast-rag index-folder ./path/to/folder --no-schema

# Check again
ast-rag evaluate --all
Expand All @@ -227,7 +227,7 @@ cypher-shell "MATCH (n) RETURN count(n)"
cypher-shell "MATCH (n:Method) RETURN count(n)"

# 3. Check indexing
grep "COMPLETE" /tmp/index_*.log | wc -l
ast-rag stats

# 4. Re-index
ast-rag init /path/to/codebase
Expand Down
Loading
Loading