Vector Search & Embeddings in Java: Building Semantic Search Engines
Vector Search & Embeddings in Java: Building Semantic Search Engines Introduction Traditional keyword-based search is dead. When users search for "best restaurants near me," they don't expect results matchin
Vector Search & Embeddings in Java: Building Semantic Search Engines
Introduction
Traditional keyword-based search is dead. When users search for "best restaurants near me," they don't expect results matching those exact wordsβthey expect restaurants that mean the right thing.
This is where vector search and embeddings enter the picture.
Vector search is the technology powering modern AI applications: ChatGPT's retrieval-augmented generation (RAG), Netflix's recommendation engine, Spotify's "Discover Weekly," and enterprise semantic search platforms. Yet many Java developers still think of search as Elasticsearch queries with exact terms.
This gap is costing you:
- Missed user intent: Returning exact matches instead of semantic matches
- Irrelevant results: Users abandon apps when search doesn't understand context
- Lost competitive advantage: Your competitors' AI-powered features work better
In this guide, you'll learn:
- What embeddings and vector search actually are (the math isn't complex, the concepts are elegant)
- How to implement semantic search in Java using production libraries
- Integration patterns with PostgreSQL (pgvector), Elasticsearch, and Pinecone
- Real-world use cases from e-commerce to support chatbots
- Performance tuning for millions of vectors
By the end, you'll understand why vector search is essential for modern applications and how to build it with Java.
Part 1: Understanding Embeddings and Vector Space
What Is an Embedding?
An embedding is a numerical representation of text, images, or other data. Instead of storing "restaurant recommendations," you store a vector of numbers: [0.25, -0.15, 0.82, ..., 0.41].
These aren't random numbers. They're learned through neural networks trained on massive datasets. Similar concepts produce similar vectors. This property is the entire foundation of vector search.
Example:
- Query: "best pizza place"
- Vector:
[0.12, 0.88, -0.31, 0.45, ...] - Document: "authentic Italian pizzeria"
- Vector:
[0.14, 0.87, -0.29, 0.46, ...]
These vectors are close in vector space. Measuring that distance (using cosine similarity, Euclidean distance, etc.) gives you a relevance score.
How Modern Embedding Models Work
In production, you don't hand-craft embeddings. You use pre-trained embedding models:
- OpenAI's text-embedding-ada-002: 1,536 dimensions, state-of-the-art for general text
- Sentence Transformers (open-source): 384-768 dimensions, fast inference, great for local deployments
- Cohere's embedding API: 4,096 dimensions, excellent domain-specific options
- Google's Gecko Embedding: Lightweight, optimized for cost
These models map text β vector in a way that preserves semantic meaning.
Vector Space Fundamentals
All vectors live in an N-dimensional space. When you have 1,536-dimensional vectors (from OpenAI), you're working in 1,536-dimensional space.
Similarity metrics:
-
Cosine Similarity (most common)
- Measures angle between vectors
- Range: -1 to 1 (1 = identical direction)
- Formula:
(A Β· B) / (||A|| Γ ||B||)
-
Euclidean Distance (L2)
- Measures straight-line distance
- Smaller = more similar
- Formula:
β(Ξ£(ai - bi)Β²)
-
Manhattan Distance (L1)
- Sum of absolute differences
- Faster but less intuitive than Euclidean
For text search, cosine similarity is almost always the right choice.
Part 2: Building Semantic Search in Java
Architecture Overview
A semantic search system has these components:
βββββββββββββββββββββββββββββββββββββββββββββββ
β User Query β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β 1. Embedding Model (Convert text β vector) β
β (OpenAI API / Local Sentence Transformer)β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β 2. Vector Search Engine β
β (Pinecone / PostgreSQL pgvector) β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β 3. Similarity Ranking β
β Return top-K most relevant results β
βββββββββββββββ¬ββββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββββββββββββββββββββ
β Results with scores (0.0 - 1.0) β
βββββββββββββββββββββββββββββββββββββββββββββββ
Example 1: Using OpenAI Embeddings + Pinecone
Step 1: Add Dependencies
<!-- pom.xml -->
<dependency>
<groupId>com.theokanning.openai-gpt3-java</groupId>
<artifactId>api</artifactId>
<version>0.18.1</version>
</dependency>
<dependency>
<groupId>io.pinecone</groupId>
<artifactId>pinecone-client</artifactId>
<version>0.1.0</version>
</dependency>
Step 2: Create Embedding Service
import com.theokanning.openai.embedding.Embedding;
import com.theokanning.openai.embedding.EmbeddingRequest;
import com.theokanning.openai.embedding.EmbeddingResult;
import com.theokanning.openai.service.OpenAiService;
public class EmbeddingService {
private final OpenAiService openAiService;
private final String modelId = "text-embedding-ada-002";
public EmbeddingService(String apiKey) {
this.openAiService = new OpenAiService(apiKey);
}
public List<Double> embedText(String text) {
EmbeddingRequest request = EmbeddingRequest.builder()
.model(modelId)
.input(Collections.singletonList(text))
.build();
EmbeddingResult result = openAiService.createEmbeddings(request);
return result.getData().get(0).getEmbedding();
}
public List<List<Double>> embedTexts(List<String> texts) {
EmbeddingRequest request = EmbeddingRequest.builder()
.model(modelId)
.input(texts)
.build();
EmbeddingResult result = openAiService.createEmbeddings(request);
return result.getData().stream()
.sorted(Comparator.comparingInt(Embedding::getIndex))
.map(Embedding::getEmbedding)
.collect(Collectors.toList());
}
}
Step 3: Pinecone Vector Search
import io.pinecone.clients.Pinecone;
import io.pinecone.clients.Index;
public class PineconeVectorStore {
private final Index index;
private final EmbeddingService embeddingService;
public PineconeVectorStore(String apiKey, String projectName,
String indexName, String embeddingApiKey) {
Pinecone client = new Pinecone.Builder()
.withApiKey(apiKey)
.build();
this.index = client.getIndex(projectName, indexName);
this.embeddingService = new EmbeddingService(embeddingApiKey);
}
// Index documents with embeddings
public void indexDocument(String docId, String content,
Map<String, String> metadata) {
List<Double> embedding = embeddingService.embedText(content);
index.upsert(
docId,
embedding,
metadata
);
}
// Search: returns top K similar documents
public List<SearchResult> search(String query, int topK) {
List<Double> queryEmbedding = embeddingService.embedText(query);
var results = index.query(queryEmbedding)
.withTopK(topK)
.withIncludeMetadata(true)
.execute();
return results.getMatches().stream()
.map(match -> new SearchResult(
match.getId(),
match.getScore(),
match.getMetadata()
))
.collect(Collectors.toList());
}
}
record SearchResult(String id, Double score, Map<String, String> metadata) {}
Step 4: End-to-End Usage
public class SemanticSearchApp {
public static void main(String[] args) {
PineconeVectorStore store = new PineconeVectorStore(
System.getenv("PINECONE_API_KEY"),
"my-project",
"restaurant-index",
System.getenv("OPENAI_API_KEY")
);
// Index documents
store.indexDocument("rest-1", "Best pizza in town, authentic Italian, family-owned",
Map.of("name", "Pasta Paradise", "type", "pizza"));
store.indexDocument("rest-2", "Fast casual ramen, Tokyo-style noodles, great broth",
Map.of("name", "Noodle House", "type", "ramen"));
// Search with semantic understanding
var results = store.search("where to find good Italian food", 3);
results.forEach(r ->
System.out.printf("ID: %s, Score: %.4f, Name: %s%n",
r.id(), r.score(), r.metadata().get("name"))
);
}
}
Example 2: Local Vector Search with PostgreSQL + pgvector
For privacy-sensitive applications, you might want embeddings to stay within your infrastructure.
Step 1: Setup PostgreSQL pgvector
-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;
-- Create table
CREATE TABLE documents (
id SERIAL PRIMARY KEY,
content TEXT NOT NULL,
embedding vector(1536),
metadata JSONB,
created_at TIMESTAMP DEFAULT NOW()
);
-- Create HNSW index for fast similarity search
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
Step 2: Java Implementation with Sentence Transformers
import org.springframework.ai.document.Document;
import org.springframework.ai.embedding.Embedding;
import org.springframework.ai.embedding.EmbeddingModel;
import org.springframework.jdbc.core.JdbcTemplate;
@Component
public class PostgresVectorStore {
private final JdbcTemplate jdbcTemplate;
private final EmbeddingModel embeddingModel;
@Autowired
public PostgresVectorStore(JdbcTemplate jdbcTemplate,
EmbeddingModel embeddingModel) {
this.jdbcTemplate = jdbcTemplate;
this.embeddingModel = embeddingModel;
}
public void indexDocument(String content, String metadata) {
Embedding embedding = embeddingModel.embed(content);
String sql = "INSERT INTO documents (content, embedding, metadata) " +
"VALUES (?, ?::vector, ?::jsonb)";
jdbcTemplate.update(sql, content, vectorToString(embedding), metadata);
}
public List<DocumentResult> similaritySearch(String query, int limit) {
Embedding queryEmbedding = embeddingModel.embed(query);
String sql = "SELECT id, content, metadata, " +
"1 - (embedding <=> ?::vector) as similarity " +
"FROM documents " +
"ORDER BY embedding <=> ?::vector " +
"LIMIT ?";
return jdbcTemplate.query(sql, new Object[]{
vectorToString(queryEmbedding),
vectorToString(queryEmbedding),
limit
}, (rs, rowNum) -> new DocumentResult(
rs.getInt("id"),
rs.getString("content"),
rs.getDouble("similarity"),
rs.getString("metadata")
));
}
private String vectorToString(Embedding embedding) {
return "[" + embedding.getOutput().stream()
.map(String::valueOf)
.collect(Collectors.joining(",")) + "]";
}
}
record DocumentResult(int id, String content, double score, String metadata) {}
Part 3: Real-World Use Cases
1. E-Commerce Product Search
Traditional: Search for "comfortable shoes" β matches only products with those exact words
Semantic: Matches "ergonomic footwear," "supportive sneakers," "cushioned athletic shoes"
public class ProductSearchService {
private final VectorStore vectorStore;
public List<Product> findSimilarProducts(String query) {
return vectorStore.search(query, 10).stream()
.map(this::toProduct)
.collect(Collectors.toList());
}
}
2. Support Chatbot with RAG
Retrieval-Augmented Generation: Embed your knowledge base, find relevant docs, feed them to LLM
@RestController
public class SupportChatbot {
private final VectorStore knowledgeBase;
private final OpenAiService openAiService;
@PostMapping("/ask")
public ResponseEntity<String> ask(@RequestBody String question) {
// Step 1: Find relevant docs using vector search
var relevantDocs = knowledgeBase.search(question, 3);
// Step 2: Build context from retrieved docs
String context = relevantDocs.stream()
.map(SearchResult::content)
.collect(Collectors.joining("\n\n"));
// Step 3: Ask LLM with context
String prompt = String.format(
"Based on this knowledge base:\n%s\n\nAnswer: %s",
context, question
);
ChatCompletionRequest request = ChatCompletionRequest.builder()
.model("gpt-4")
.messages(List.of(new ChatMessage(ChatMessageRole.USER.value(), prompt)))
.build();
String answer = openAiService.createChatCompletion(request)
.getChoices().get(0).getMessage().getContent();
return ResponseEntity.ok(answer);
}
}
3. Duplicate Detection
Find near-duplicate documents or suspicious fraud patterns
public class DuplicateDetector {
private final VectorStore vectorStore;
public boolean isProbablyDuplicate(String document, double threshold) {
var similarDocs = vectorStore.search(document, 1);
return !similarDocs.isEmpty() &&
similarDocs.get(0).score() > threshold;
}
}
Part 4: Performance Optimization
Batching Embeddings
Don't embed one document at a time. Batch them.
public void indexManyDocuments(List<Document> documents) {
// β Slow: N API calls
// documents.forEach(doc -> index(doc));
// β
Fast: 1 API call per batch
Iterables.partition(documents, 100).forEach(batch -> {
List<String> texts = batch.stream()
.map(Document::getContent)
.collect(Collectors.toList());
List<List<Double>> embeddings = embeddingService.embedTexts(texts);
for (int i = 0; i < batch.size(); i++) {
vectorStore.index(batch.get(i).getId(), embeddings.get(i));
}
});
}
Caching Embeddings
Store computed embeddings to avoid redundant API calls
@Component
public class CachedEmbeddingService {
private final EmbeddingService service;
private final Map<String, List<Double>> cache = new ConcurrentHashMap<>();
public List<Double> embed(String text) {
return cache.computeIfAbsent(text, key -> service.embedText(key));
}
}
Dimension Reduction
Not all 1,536 dimensions are necessary for your use case. PCA can reduce them:
public List<Double> reduceDimensions(List<Double> embedding, int targetDim) {
// Use Apache Commons Math or similar
// This trades accuracy for speed/storage
return PCA.reduce(embedding, targetDim);
}
Part 5: Production Considerations
1. Latency Budget
- Embedding API call: 50-200ms
- Vector search: 10-50ms (with proper indexing)
- Total for user query: <500ms
If this is too slow, cache frequent queries.
2. Cost Management
- OpenAI embeddings: ~$0.02 per million tokens
- Pinecone: ~$0.70 per 1M vectors/month
- Self-hosted (pgvector): Your infrastructure cost
For high-volume applications, self-hosting saves money.
3. Handling Updates
When document content changes, re-embed and update:
public void updateDocument(String docId, String newContent) {
List<Double> newEmbedding = embeddingService.embedText(newContent);
vectorStore.update(docId, newEmbedding, newContent);
}
4. Monitoring
Track embedding quality:
@Component
public class EmbeddingQualityMonitor {
private final MeterRegistry meterRegistry;
public void recordSimilarityScore(double score) {
Timer.builder("search.similarity.score")
.publishPercentiles(0.5, 0.95, 0.99)
.register(meterRegistry)
.record(Duration.ofMillis((long) (score * 1000)));
}
}
Conclusion
Vector search and embeddings are no longer bleeding-edge. They're essential for:
- Better search relevance (understand user intent, not just keywords)
- Semantic recommendations (similarity-based, not just collaborative filtering)
- RAG-powered chatbots (ground LLMs in your data)
- Anomaly detection (spot unusual patterns in vector space)
Key Takeaways:
- Embeddings are numerical representations where similar concepts = close vectors
- Use cosine similarity to measure relevance
- Choose between managed (Pinecone, Weaviate) or self-hosted (PostgreSQL + pgvector)
- Batch API calls, cache embeddings, monitor quality
- Start with OpenAI embeddings; optimize to self-hosted later if cost-prohibitive
Your next semantic search implementation is just these components away. Build it today.
Further Reading
- Pinecone Docs: https://docs.pinecone.io/
- Sentence Transformers: https://www.sbert.net/
- pgvector Extension: https://github.com/pgvector/pgvector
- RAG Best Practices: https://docs.anthropic.com/en/api/guides/vision
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.