Vector Databases / Векторные базы данных
Год издания: 2026
Автор: Borwankar Nitin / Борванкар Нитин
Издательство: O’Reilly Media, Inc.
ISBN: 978-1-098-17759-1
Язык: Английский
Формат: PDF/EPUB
Качество: Издательский макет или текст (eBook)
Интерактивное оглавление: Да
Количество страниц: 293
Описание: The AI revolution is here, and at its core lies a game-changing technology that most developers haven’t fully explored: vector databases. From powering semantic search to enabling large language models (LLMs) and generative AI, vector databases are reshaping how we build applications with unstructured data like text, images, and audio. But how do you go from curious to capable with this vital technology? That’s where this book comes in.
In this hands-on guide, author Nitin Borwankar takes you through the “why, what, and how” of vector databases, starting with the basic theory behind vector embeddings and progressing to building applications with real-world tools. You’ll learn about Word2vec, how to convert open source SQL databases like SQLite3 and PostgreSQL into vector databases, and integrate them into retrieval-augmented generation (RAG) applications. Whether you’re a Python developer, data engineer, or ML practitioner, this book gives you the foundation to leverage vector databases confidently in your AI projects.
Understand the connection between vector databases, embeddings, and LLMs
Learn practical approaches for transforming SQL databases into vector databases
Build RAG applications for both personal and enterprise use
Apply vector databases to solve real-world AI challenges
Learn how to use vector databases with LLMs to build applications
Революция в области искусственного интеллекта уже началась, и в ее основе лежит технология, меняющая правила игры, которую большинство разработчиков еще не изучили до конца: векторные базы данных. Начиная с семантического поиска и заканчивая большими языковыми моделями (LLM) и генеративным искусственным интеллектом, векторные базы данных меняют наши подходы к созданию приложений с неструктурированными данными, такими как текст, изображения и аудио. Но как с помощью этой жизненно важной технологии перейти от любопытства к способностям? Вот тут-то и пригодится эта книга.
В этом практическом руководстве автор Нитин Боруанкар расскажет вам о том, “почему, что и как” работает с векторными базами данных, начиная с базовой теории, лежащей в основе векторных встраиваний, и переходя к созданию приложений с использованием реальных инструментов. Вы узнаете о Word2vec, о том, как преобразовать базы данных SQL с открытым исходным кодом, такие как SQLite3 и PostgreSQL, в векторные базы данных и интегрировать их в приложения с расширенным поиском (RAG). Независимо от того, являетесь ли вы разработчиком Python, инженером по обработке данных или специалистом по ML, эта книга даст вам основу для уверенного использования векторных баз данных в ваших проектах искусственного интеллекта.
Поймите связь между векторными базами данных, внедрениями и LLMS
Изучите практические подходы к преобразованию баз данных SQL в векторные базы данных
Создайте приложения RAG как для личного, так и для корпоративного использования
Примените векторные базы данных для решения реальных задач искусственного интеллекта
Узнайте, как использовать векторные базы данных с LLMS для создания приложений
Примеры страниц (скриншоты)
Оглавление
Preface xi
1. Introduction to Vector Databases 1
Why Do You Need Vector Databases? 1
A New Data Type: Vector 2
Similarity Search 3
What’s Different About the Vector Type? 6
Where Do You Use Vector Databases? 8
SQL Versus Vector Databases 9
The Foundation of Business Math: Accounting Arithmetic 9
Vector Representation in a Relational Database Management System 10
The Need for Vector-Specific Capabilities 11
NoSQL Versus Vector Databases 11
NoSQL Databases and Vector Storage 11
Limitations of Vector Extensions in NoSQL Databases 12
When to Choose NoSQL with Vector Extensions 12
Hybrid Approaches: Combining Structured and Vector Data 13
The Need for Both Vector Data and Metadata 13
Limitations of Pure Vector Storage 13
Hybrid Database Architecture 14
Example of a Hybrid Query 14
Benefits of the Hybrid Approach 15
Conclusion 15
2. Embeddings 17
Understanding Vector Embeddings: Why We Need Them 17
Word2Vec: The Breakthrough That Changed Everything 19
Doc2Vec: From Words to Documents 20
From Embeddings to Modern Language Models:
The Transformer Connection 23
Encoder-Only Transformers (BERT and Its Variants) 24
Decoder-Only Transformers (GPT Family) 24
Encoder-Decoder Transformers (T5, BART) 25
Embedding Models: The Specialized Vector Generators 26
Distinction from Traditional Models 26
Role in Modern LLM Applications 27
Practical Applications and Use Cases 28
Simple RAG Pipeline 28
The sentence-transformers Library: The Swiss Army Knife
of Text Embeddings 31
Best Practices for Using SentenceTransformers: A Detailed Guide 35
The Embedding Layer: The Gateway to Zero-Shot Learning 40
Anatomy of Transformer Embeddings 40
Connection to Zero-Shot Learning 42
Key Characteristics That Enable Zero-Shot Learning 43
Limitations and Considerations 45
Latest Developments and Trends 46
Vector Arithmetic with Word2Vec: A Hands-On Guide 46
Step 1: Setup and Installation 46
Step 2: Load Pretrained Word2Vec Model 46
Step 3: Implement Vector Arithmetic Functions 47
Step 4: Classic King–Queen Analogy 48
Step 5: More Interesting Analogies 49
Step 6: Interactive Exploration Tool 49
Final Words on Vector Arithmetic 50
Conclusion 51
3. Similarity Search with FAISS 53
Foundations 53
Vector Representations 55
Distance Metrics 56
Selection Heuristics 58
FAISS Indexes 58
Flat Indexes (Brute Force) 58
IVF-Based Indexes 59
LSH-Based Indexes 60
HNSW-Based Indexes 61
Other Specialized Indexes 61
Composite and Transformative Indexes 62
Choosing the Right Index 62
Quantization 65
SQ 65
PQ 67
The ANN Problem 71
The Problem 72
Avoid Computational Cost 72
Key ANN Techniques in FAISS 73
Choosing an Index in FAISS 75
Code Example 75
Understanding HNSW Indexes 76
What Is HNSW? 77
How HNSW Works 78
Key Parameters Explained 79
Practical Example: Building a Similarity Search System 80
Performance Characteristics 81
Best Practices 82
FAISS Architecture and Components 83
Foundation 83
Core Concepts 85
Key Components 85
Common Workflow 87
Illustrative Example 87
Key Takeaways 88
Further Exploration 88
Conclusion 89
4. Semantic Search with SQLite3 91
Understanding the SQLite Vector Similarity Search Extension 91
Core Capabilities 92
Architecture Overview 93
Limitations 93
Setting Up the Development Environment 94
Installing Dependencies 94
Verifying the Installation 95
Operational Pragmas 96
Designing the Database Schema 96
Schema Requirements 96
Table Definitions 96
Schema Design Decisions 98
Connecting to Reddit with the Python Reddit API Wrapper 99
Creating Reddit API Credentials 99
PRAW Client Implementation 99
Usage Example 102
Content Extraction and Preprocessing 102
Text Cleaning Pipeline 102
Quality Filtering 105
Generating and Storing Embeddings 105
Embedding Generator 106
Database Storage 108
Batch Processing Pipeline 112
Building the Vector Index 113
Understanding VSS Indexing 113
Index Management 114
Implementing Semantic Search 117
Search Result Container 117
Search Engine 117
Putting It All Together 123
Workflow Example 124
Example Output 126
Extension: Incremental Indexing 127
Conclusion 129
5. Building an ArXiv Paper Search System with PostgreSQL pgvecto 131
The Challenge of Searching Scientific Literature 131
Why ArXiv Makes an Ideal Data Source 131
Real-World Use Cases 132
Technology Stack Rationale 132
Architecture Overview 133
System Components 133
Data Flow 134
Design Philosophy 135
Environment Setup and Dependencies 136
PostgreSQL and pgvector Installation 136
Python Environment Configuration 137
Directory Structure and Configuration 137
Verification and Testing 138
Database Design for Scientific Papers 139
Schema Design Principles 139
Core Tables Structure 140
Vector Storage Strategy 144
Indexing Strategy 144
ArXiv Integration and PDF Management 145
ArXiv API Client Implementation 146
PDF Download Pipeline 147
Batch Processing System 148
PDF Text Extraction and Processing 150
PDF Extraction Challenges 150
Intelligent Text Chunking 152
Embedding Generation and Storage 153
Embedding Model Strategy 154
Batch Processing Pipeline 155
Similarity Search Implementation 156
Interactive Application and UI 158
Docker Packaging for Local Deployment 160
Container Architecture 161
Docker Compose Configuration 161
Database Initialization Scripts 163
Development Workflow 164
Cloud-Ready Design 164
Basic Performance Tuning 165
Index Configuration 165
Query Performance 165
Resource Management 165
Next Steps 165
Current Limitations 165
Enhancement Ideas 166
What We Did 166
System Achievements 166
Technical Skills Gained 167
Practical Research Tool 167
Foundation for Advanced Systems 167
Future Potential 167
Conclusion 167
6. Building a Retrieval-Augmented Generation System with SQLite VSS and Ollama 169
System Architecture Overview 170
Database Foundation with Vector Support 171
Setting Up the Vector-Enabled Database 171
Schema Design for RAG 172
Creating Search Indexes 173
Text Processing and Embedding Generation 174
Embedding Model Management 174
Intelligent Text Chunking 175
Storing Content with Embeddings 176
Hybrid Search Implementation 177
Hybrid Search Algorithm 177
Semantic Search Component 178
Keyword Search Component 179
Score Fusion and Ranking 180
LLM Integration with Ollama 181
Ollama API Client 181
Health Check Function 182
The RAG Pipeline 182
Context Formatting 182
Question-Answering Pipeline 183
Demonstration and Testing 185
Sample Data Loading 185
Main Demonstration Function 186
Interactive Q&A Interface 187
Quick Testing Utility 188
Next Steps: Extending the System 188
Missing Reddit Data Features 188
Performance Optimizations 190
Production Considerations 190
Advanced RAG Patterns 191
Conclusion 191
7. Building a Scientific RAG System with PostgreSQL and pgvector 193
System Goals and Capabilities 194
Architecture Overview 194
Database Foundation with pgvector 196
Database Configuration and Setup 197
Schema Design for Scientific Papers 197
High-Performance Vector Indexes 199
Embedding Generation Strategy 199
ArXiv Integration and PDF Processing 200
Paper Discovery with ArXiv API 200
Intelligent PDF Text Extraction 201
Advanced Text Chunking 203
Storage Pipeline with Embeddings 204
Multilevel Semantic Search 206
Abstract-Level Search 206
Section-Level Search 207
The RAG Pipeline: Deep Dive 208
Local LLM Integration with Ollama 209
Health Check and Model Discovery 210
Intelligent Context Retrieval 210
Scientific Prompt Engineering 211
Complete RAG Execution Pipeline 212
Demonstration and Interactive Interface 213
Main Demonstration Flow 213
Search Demonstrations 214
RAG Demonstration 215
Interactive Search Interface 216
Entry Point with Mode Selection 217
Technical Note on HNSW 217
How to Evaluate Your Results 219
Next Steps: Extending the Scientific RAG System 220
Conclusion 223
8. Building a Complete Conversation Search and RAG System 225
System Goals and Capabilities 226
System Architecture Overview 227
What We’ll Build Together 229
Database Foundation for Conversation Storage 230
Designing the Conversation Schema 230
Three-Table Architecture for Optimal Performance 231
High-Performance Vector Indexing 232
Conversation Import and Data Processing Pipeline 233
Robust JSON Import with Error Handling 233
Atomic Transaction Processing 234
Timestamp Handling and Data Validation 234
Error Recovery and Logging 235
Efficient Embedding Generation and Batch Processing 236
Singleton Pattern for Model Management 236
Incremental Processing Strategy 237
Batch Processing for Optimal Performance 237
Database Insertion with Conflict Handling 238
Contextual Search with Conversational Understanding 239
Semantic Similarity Search 239
Multitable Joins for Rich Context 240
Result Formatting and Structure 240
Conversation Context Retrieval 241
Context Window Calculation 241
RAG Integration for Conversation History 242
Structured Context Management 242
Local LLM Integration with Ollama 243
Health Monitoring and Model Discovery 243
Context Retrieval and Assembly 244
Conversational Prompt Engineering 244
Complete RAG Pipeline with Performance Monitoring 245
Complete Web API with FastAPI 247
FastAPI Application Structure 247
Request Models with Validation 247
Search Endpoint Implementation 248
RAG Question-Answering Endpoint 248
System Statistics and Monitoring 248
Server Startup and Configuration 249
Demonstration and Sample Data 250
Realistic Sample Data Generation 250
Multitopic Sample Coverage 251
Sample Data Processing Pipeline 252
Comprehensive System Demonstration 252
Progressive Feature Demonstration 253
RAG Demonstration with Conditional Execution 254
Production Import Functionality 254
Application Entry Points 255
Conclusion: A Complete Personal Knowledge System 255
9. Vector Query Language 257
Core Concepts 258
Data Model 258
Basic Syntax Structure 259
Vector Operations 260
Similarity Search 260
Hybrid Search 261
Range Search 261
Batch Operations 261
Vector Functions and Aggregations 262
Vector Functions 262
Vector Aggregations 263
Index 265