rag-architect
>
Works with
Claude CodeCursorCodex CLIGitHub CopilotGemini CLI
--- name: rag-architect description: > license: MIT --- # RAG Architect Design retrieval-augmented generation pipelines with the right tradeoffs at each layer. ## Workflow 1. **Choose chunking strategy** -- Match chunk method to document structure 2. **Select embedding model** -- Balance dimensions, speed, and domain fit 3. **Choose vector DB** -- Match scale, features, and deployment model 4. **Design retrieval** -- Dense, sparse, hybrid, or reranked 5. **Evaluate** -- Measure faithfulness, relevance, and answer quality ## Reading Guide | Decision | File | | ---------------------------------------------- | ------------------------------------------------------------ | | Chunking strategies + embedding models | [chunking-and-embedding.md](./chunking-and-embedding.md) | | Retrieval strategies + vector DBs + evaluation | [retrieval-and-evaluation.md](./retrieval-and-evaluation.md) | ## Quick Decision Matrix | Document type | Chunking | Embedding | Retrieval | | -------------- | ------------------------- | ---------------- | --------------- | | Code | Semantic (AST-aware) | Code-specialized | Hybrid + rerank | | Legal/medical | Document-aware (sections) | Domain-specific | Dense + rerank | | Chat logs | Sentence | General-purpose | Dense | | Technical docs | Recursive | General-purpose | Hybrid | | Mixed/unknown | Recursive (fallback) | General-purpose | Hybrid + rerank | ## What You Get - A RAG pipeline architecture specifying chunking strategy, embedding model, vector store, and retrieval method for your document types - Concrete configuration recommendations (chunk size, overlap, dimensions, top-k) with rationale for each tradeoff - An evaluation plan using RAGAS or equivalent metrics to validate retrieval quality before and after tuning ## Rules 1. Start simple -- fixed-size chunks + dense retrieval is a valid baseline 2. Measure before optimizing -- run RAGAS evaluation before adding complexity 3. Chunk overlap matters -- 10-20% overlap prevents context loss at boundaries 4. Embedding dimensions are a tradeoff -- higher is not always better (cost, latency) 5. Hybrid retrieval (dense + sparse) almost always beats either alone
More AI & ML skills
writing-shape
mattpocock/skills
Writing, exploit: shape raw material into an article, paragraph by paragraph.
279.3k
writing-fragments
mattpocock/skills
Writing, explore: mine raw fragments, no structure yet.
279.2k
full-output-enforcement
leonxlnx/taste-skill
Overrides default LLM truncation behavior. Enforces complete code generation, bans placeholder patterns, and handles token-limit splits cleanly. Apply to any task requiring exhaustive, unabridged output.
275.4k

