Skip to content

About

Intel® AI for Enterprise RAG converts enterprise data into actionable insights with excellent TCO. Utilizing Intel Xeon processors ensuring streamlined deployment.

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Intel® AI for Enterprise RAG

License: Apache 2.0 Component of: AI Solutions Platform: Intel Xeon Pipelines: ChatQnA · DocSum · AudioQnA Identity: Keycloak

Enterprise-grade retrieval-augmented generation layer for Intel® AI for Enterprise Solutions. Turn your enterprise documents into a governed, production-ready AI assistant on Intel® Xeon® CPUs.

Ships the Ansible roles, Helm charts, and composable pipeline definitions that deploy a full RAG application - document ingestion, vector search, reranking, guardrails, chat history, and a web UI - wired into the platform's identity, gateway, storage, and observability.

Important

This repository is not used standalone. It is a component of the ai-solutions platform and is automatically cloned into it at enterprise-ai-solutions/ext/enterprise.ai-erag/, where it contributes the erag layer. Install the platform first - it provisions the Kubernetes cluster, the platform services (cert-manager, Istio, MetalLB, Envoy Gateway, PostgreSQL, Keycloak, MinIO, observability), and the model serving this layer consumes. Every command below runs from the solutions repo root, not from here.


What is Intel® AI for Enterprise RAG?

Intel® AI for Enterprise RAG deploys that whole path as one opt-in layer. You pick a pipeline flavour, run two commands, and get a working assistant grounded in your own documents - no model training or fine-tuning required.

Pipelines are composed, not hardcoded. A flavour declares an ordered flow of steps, the composer renders it into a GMConnector resource, and the GMC operator reconciles the microservices behind it. Swapping a retrieval strategy or adding output guardrails is a config change, not a rewrite.

Want the full picture? See Architecture and Pipelines.

What makes it enterprise-grade

  • Access control, not just an API key - Keycloak OIDC single sign-on across the UI, Grafana and Keycloak itself, a guardrail on every query by default, and opt-in role-based access control that scopes retrieval to the documents each user is cleared to see.
  • Modular pipelines, not a monolith - a flavour declares an ordered flow of steps, the composer renders it into a GMConnector resource, and the operator reconciles the microservices behind it. Changing retrieval strategy or adding output guardrails is a config edit, not a rewrite.
  • Integrated, not bolted on - identity, TLS, object storage, PostgreSQL and Grafana come from the shared platform, so RAG becomes one more governed workload instead of a second stack to operate.
  • Secure by default - Pod Security Standards enforcement, Istio ambient mTLS between services, and generated credentials with no secrets in the repo.
  • Four workloads, one stack - conversational retrieval (ChatQnA), document summarization (DocSum), voice question answering (AudioQnA), and translation.
  • Tuned for Intel® Xeon® - horizontal pod autoscaling and NUMA-aware CPU pinning through the platform's balloons policy.

Architecture

A question enters through the gateway, is authenticated against Keycloak, and reaches the pipeline router. The router walks the composed flow - embed the query, retrieve candidates from the vector database, rerank them, apply input guardrails, build the prompt, and call the LLM - then streams the grounded answer back. Model inference itself is served by the platform's inference layer, so the RAG namespace runs no model servers of its own.

Intel AI for Enterprise RAG ChatQnA architecture: a request from the UI enters the gateway, is authenticated, and passes through embedding, retrieval from the vector database, reranking, guardrails, and prompt templating before reaching the LLM, with document ingestion feeding the vector database from the object store

DocSum's flow is shown in architecture_docsum.svg, and the full microservice map in microservices_architecture.png.

See the Architecture reference for the component inventory, the roles that deploy them, and how this repo plugs into the platform.


Quick Start

Deploys the RAG stack on a single node with defaults. In three steps you will have an assistant answering questions about documents you ingest.

Note

Prerequisites: a Xeon host with 60 logical cores, 128 GB RAM, and 200 GB free disk; Ubuntu 22.04/24.04; passwordless sudo; internet access; and Hugging Face access to the default models. Full list, including the 32-core limited deployment → Prerequisites.

Step 1 - Install the stack

init erag clones this repo and the inference layer it depends on, at the revisions the platform pins, and seeds the configuration for your chosen pipeline. install erag then pulls in the layers below it (infrastructure, platform, inference) and deploys the RAG application on top.

git clone https://github.com/intel/enterprise-ai-solutions.git
cd enterprise-ai-solutions

./es_auto_installer.sh configure     # one-time machine prep (Python 3.11+, yq, kubectl, helm)
./es_auto_installer.sh init erag     # clone + seed the erag layer and its dependencies
./es_auto_installer.sh install erag  # deploy everything

Settings for this layer land in env/local/config.erag.yaml. Pick a different pipeline with --flavour:

./es_auto_installer.sh init erag --flavour docsum

Note

For any parameters, customization options or multinode deployment, refer to documentation.

Tip

--env defaults to local. Tear down with ./es_auto_installer.sh teardown erag. Install and teardown are environment-scoped: if you installed with --env prod, you must tear down with --env prod.

Step 2 - Reach the UI

The gateway binds ports 80 and 443 on the node, so no port forwarding is needed. Add each subdomain to /etc/hosts on the machine you browse from - wildcards do not work there:

<node-ip> solutions.ai grafana.solutions.ai keycloak.solutions.ai s3.solutions.ai seaweedfs.solutions.ai

Then open https://solutions.ai. First-login credentials are written to env/local/logs/rag/default_credentials.txt; you will be asked to change the password immediately.

Important

With the default self-signed certificates, visit https://s3.solutions.ai once and accept the warning before ingesting documents. Not needed with custom certificates.

Step 3 - Ingest a document and ask about it

Sign in as the admin user, open the Admin Panel → Data Ingestion tab, and upload a file or point it at a URL. Once ingestion reports complete, ask a question in the chat and the answer will cite your document.

To verify the pipeline from the command line instead:

cd ext/enterprise.ai-erag/deployment
./scripts/test_connection.sh    # ChatQnA; use test_docsum.sh or test_translation.sh for those flavours

What you get

Area Component What it gives you
Ingestion Enhanced Data Preparation (EDP) Extract, split, and embed documents from the object store or SharePoint, with opt-in scheduled sync
Retrieval Vector database + reranking Redis Cluster, PGVector, or Microsoft SQL Server backends, with opt-in per-user access control on the index
Pipelines GMC operator + composer ChatQnA, DocSum, AudioQnA, and translation flows, composed from shared steps
Safety Guardrail microservices Query filtering on by default; response filtering via the output_guard variant, ingestion filtering via edp_dp_guard_enabled
Access Keycloak OIDC + APISIX Single sign-on, realm roles per persona, and opt-in per-user document scoping
Agents MCP gateway Expose retrieval and ingestion to AI agents over Model Context Protocol

Curious first? The demo below shows ChatQnA in action.

Note

The video showcases an earlier release. The current UI, installation flow, and feature set have moved on since it was recorded.


Advanced

The Quick Start deploys the ChatQnA flavour with defaults. From here you can change the pipeline and its variants, swap models, enable multilingual retrieval, connect SharePoint or an external S3 store, wire up agents, and tune every microservice.

Goal Guide
Deploy step by step, endpoints and credentials Deploy the RAG layer
Pipelines, flavours, and variants Pipelines
Every configuration option Configuration
Change the LLM, embedding, or reranking model Models
SSO, MFA, and Active Directory federation Authentication
Ingest from SharePoint Online SharePoint
Connect AI agents MCP Integration
External S3 or NetApp ONTAP document store Object Store
Components, roles, and request flow Architecture
Dashboards and logs Telemetry
Scaling and tuning Performance
Something isn't working Troubleshooting
VMware deployment Deploy on VMware
What it is and why it exists Meet Intel® AI for Enterprise RAG
Common questions FAQ
Terminology Glossary

Publications


Support

Submit questions, feature requests, and bug reports on the GitHub Issues page.

License

Intel® AI for Enterprise RAG is licensed under the Apache License Version 2.0. Refer to the "LICENSE" file for the full license text and copyright notice.

This distribution includes third-party software governed by separate license terms. This third-party software, even if included with the distribution of the Intel software, may be governed by separate license terms, including without limitation, third-party license terms, other Intel software license terms, and open-source software license terms. These separate license terms govern your use of the third-party programs as set forth in the "THIRD-PARTY-PROGRAMS" file.

Please note: component(s) depend on software subject to non-open source licenses. If you use or redistribute this software, it is your sole responsibility to ensure compliance with such licenses.

Security

The Security Policy outlines our guidelines and procedures for ensuring the highest level of security and trust for our users who consume Intel® AI for Enterprise RAG.

Intel's Human Rights Principles

Intel is committed to respecting human rights and avoiding complicity in human rights abuses. See Intel's Global Human Rights Principles. Intel's products and software are intended only to be used in applications that do not cause or contribute to a violation of an internationally recognized human right.

Model Card Guidance

You, not Intel, are responsible for determining model suitability for your use case. For information regarding model limitations, safety considerations, biases, or other information consult the model cards (if any) for models you use, typically found in the repository where the model is available for download. Contact the model provider with questions. Intel does not provide model cards for third party models.

Contributing

If you want to contribute to the project, please refer to the guide in CONTRIBUTING.md file.

Links


Intel, the Intel logo, OpenVINO, the OpenVINO logo, Pentium, and Xeon are trademarks of Intel Corporation or its subsidiaries. Other names and brands may be claimed as the property of others.

© Intel Corporation

About

Intel® AI for Enterprise RAG converts enterprise data into actionable insights with excellent TCO. Utilizing Intel Xeon processors ensuring streamlined deployment.

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages