Elastic Runs Jina AI Models Fully On-Premise

Elastic has announced a new offering that provides enterprise-grade data extraction and semantic search with no internet connection or third-party AI services required. It is aimed at regulated industries, air-gapped environments and teams wanting full control over data, cost and performance.

Organisations in regulated or disconnected environments can now run Jina AI models entirely on their own infrastructure through Jina On-Prem, released by Elastic.

Jina AI is a Berlin-based company founded in 2020 and known for its open-source, search-grade embedding and reranking models, used widely for neural search and retrieval-augmented generation. Elastic acquired Jina AI in October 2025, and Jina On-Prem is the first major packaging of those models for the Elastic stack.

The underlying models power the machinery beneath modern search. They turn documents into embeddings, which are numerical representations of meaning, then rerank results for relevance and read and extract text from files. Semantic search built on embeddings retrieves records by concept rather than exact keywords, so a query can surface a relevant document even when it shares no words with the search terms.

Jina On-Prem covers text, images, audio and video in a single embedding space, running within a customer's environment and making no outbound calls once deployed. Small Jina models run on a single 8GB GPU while matching the accuracy of much larger models that need many times more GPU memory.

The suite includes all 28 Jina AI models, among them the jina-embeddings-v5-omni multimodal model and the jina-reranker-v3 reranking model. It installs with a single command and makes no licence-server, telemetry or model-registry calls. It supports both CPU and GPU hardware with automatic GPU detection, and applications reach it through standard API schemas so existing integrations work without rewriting code.

"Historically, teams running search and retrieval in regulated or disconnected environments have had to choose between capability and control," said Ajay Nair, general manager for Elasticsearch and Platform at Elastic. He said running Jina models fully on-premises removed that compromise.

Many Australian and New Zealand government agencies run OpenText Content Manager as their system of record while using Elastic as a cross-repository enterprise search layer over records, network drives and email. Jina On-Prem drops into that existing Elastic tier as a replacement for the models served through Elastic Inference Service, adding semantic, meaning-based retrieval across those repositories without re-architecting.

The decisive feature for that audience is data sovereignty. Records classified at PROTECTED and above cannot be sent to external AI services, which has kept many agencies away from semantic search and retrieval-augmented generation. Running the models on-premises with no outbound calls lets an agency improve the findability of its records, and build AI assistants grounded in them, while staying inside its security boundary and IRAP obligations.

Beyond search, the same models underpin retrieval-augmented generation, where an AI assistant grounds its answers in an organisation's own documents. Running that pipeline on-premises means the grounding data never leaves the environment. Elastic said the small Jina models match the accuracy of far larger ones while running at a fraction of the compute cost, lowering the hardware barrier for agencies that cannot rely on cloud GPUs.

The release also lets air-gapped Elastic deployments adopt the models without changing how their applications call embedding and reranking services, easing the path for sites already invested in the Elastic stack.

Jina On-Prem is available now for download via GitHub through an access token, with installation instructions on the Jina On-Prem Quick Start page.

www.elastic.co 

 

Business Solution