NEW · OFFICIAL PYTHON PACKAGE

RAG on managed Apache Solr — one pip install

Opensolr is now a native LangChain vector store. Server-side GPU embeddings, hybrid BM25 + kNN search, and a managed Solr 9 index behind every retriever — with zero embedding models to configure.

$ pip install langchain-opensolr
Python 3.9+ LangChain 0.3+ & 1.x Apache Solr 9.x MIT licensed

The whole tutorial

This is not a teaser — this is the entire integration. No embedding model, no API key juggling, no schema design.

from langchain_opensolr import OpensolrVectorStore

vs = OpensolrVectorStore(
    index="mysite__dense",          # vector-enabled Opensolr index
    email="you@example.com",
    api_key="YOUR_OPENSOLR_API_KEY",
    create_if_missing=True,          # provisions the index on first use
)

vs.add_texts(["Hybrid search fuses BM25 with vector similarity",
              "Cats sleep sixteen hours a day"])

docs = vs.similarity_search("how do lexical and semantic search combine?")

No embedding model was configured, because embedding happens server-side — on Opensolr's GPU infrastructure, at both index and query time.

What happens under the hood

Your app talks to one package. The package talks to a managed pipeline that already exists.

Your Python app chains · agents · RAG as_retriever() langchain-opensolr OpensolrVectorStore OpensolrEmbeddings auto host + auth discovery GPU Embeddings E5-large-instruct · multilingual 1024 dims · cosine api.opensolr.com Managed Solr 9 BM25 + kNN · {!hybrid} facets · filters · highlighting us · de · fi Ranked Documents with scores + metadata

Why this is different

Most vector stores in the LangChain directory make you bring your own embedding model and run your own database. This one doesn't.

Zero embedding config

The shortest constructor in the directory: api_key + index. Texts and queries are embedded server-side on GPU — no OpenAI key, no local model, no sentence-transformers install.

True hybrid search

Pure vector search fails on exact identifiers; pure BM25 fails on meaning. hybrid=True fuses both scores per document, with four modes and a tunable alpha balance.

Lossless metadata + filters

Metadata round-trips exactly as you stored it, and filters work both ways: filter={"category": "docs"} or any raw Solr fq expression when you need real power.

Auto index provisioning

create_if_missing=True provisions a vector-enabled Solr 9 index on first use — pick us, de or fi. No servers, no schema files, no ZooKeeper.

Drops into any chain

vs.as_retriever() and it's a standard LangChain retriever — RAG tutorials, agents, LangGraph, LCEL pipelines. Everything that accepts a retriever accepts this.

Plain Apache Solr underneath

Every index is also a real Solr core with the native /select API — facets, highlighting, spellcheck, stats. When you outgrow the vector-store interface, nothing is locked away.

Two lines you'll actually use

Hybrid retrieval with metadata filters, and the retriever that plugs into every RAG example ever written.

Hybrid search, filtered

docs = vs.similarity_search(
    "affordable restaurants",
    k=5,
    hybrid=True,      # BM25 + kNN, fused
    mode="union",     # or keywords_required,
                      # meaning_required,
                      # intersection
    alpha=0.5,        # semantic ↔ lexical
    filter={"city": "Cluj"},
)

As a retriever, in any chain

retriever = vs.as_retriever(
    search_kwargs={"k": 5, "hybrid": True}
)

# now use it anywhere LangChain
# expects a retriever:
chain = (
    {"context": retriever, "question": ...}
    | prompt | llm
)

Build your first RAG pipeline tonight

Free 15-day trial, no credit card. The included AI quota comfortably covers the whole tutorial — index, embed, search, and retrieve.

Germany · Solr 9.0 Finland · Solr 9.6 USA · Solr 9.6

Need a vector-enabled environment in another region? We deploy them on request — dedicated, in the region you choose (paid add-on).