Enterprise-grade retrieval-augmented generation layer for Intel® AI for Enterprise Solutions. Turn your enterprise documents into a governed, production-ready AI assistant on Intel® Xeon® CPUs.
Ships the Ansible roles, Helm charts, and composable pipeline definitions that deploy a full RAG application - document ingestion, vector search, reranking, guardrails, chat history, and a web UI - wired into the platform's identity, gateway, storage, and observability.
Important
This repository is not used standalone. It is a component of the
ai-solutions
platform and is automatically cloned into it at enterprise-ai-solutions/ext/enterprise.ai-erag/, where it contributes the
erag layer. Install the platform first - it provisions the Kubernetes cluster,
the platform services (cert-manager, Istio, MetalLB, Envoy Gateway, PostgreSQL, Keycloak,
MinIO, observability), and the model serving this layer consumes.
Every command below runs from the solutions repo root, not from here.
Intel® AI for Enterprise RAG deploys that whole path as one opt-in layer. You pick a pipeline flavour, run two commands, and get a working assistant grounded in your own documents - no model training or fine-tuning required.
Pipelines are composed, not hardcoded. A flavour declares an ordered flow of steps, the composer renders it into a GMConnector resource, and the GMC operator reconciles the microservices behind it. Swapping a retrieval strategy or adding output guardrails is a config change, not a rewrite.
Want the full picture? See Architecture and Pipelines.
- Access control, not just an API key - Keycloak OIDC single sign-on across the UI, Grafana and Keycloak itself, a guardrail on every query by default, and opt-in role-based access control that scopes retrieval to the documents each user is cleared to see.
- Modular pipelines, not a monolith - a flavour declares an ordered flow of steps, the composer renders it into a
GMConnectorresource, and the operator reconciles the microservices behind it. Changing retrieval strategy or adding output guardrails is a config edit, not a rewrite. - Integrated, not bolted on - identity, TLS, object storage, PostgreSQL and Grafana come from the shared platform, so RAG becomes one more governed workload instead of a second stack to operate.
- Secure by default - Pod Security Standards enforcement, Istio ambient mTLS between services, and generated credentials with no secrets in the repo.
- Four workloads, one stack - conversational retrieval (ChatQnA), document summarization (DocSum), voice question answering (AudioQnA), and translation.
- Tuned for Intel® Xeon® - horizontal pod autoscaling and NUMA-aware CPU pinning through the platform's balloons policy.
A question enters through the gateway, is authenticated against Keycloak, and reaches the pipeline router. The router walks the composed flow - embed the query, retrieve candidates from the vector database, rerank them, apply input guardrails, build the prompt, and call the LLM - then streams the grounded answer back. Model inference itself is served by the platform's inference layer, so the RAG namespace runs no model servers of its own.
DocSum's flow is shown in architecture_docsum.svg, and the full microservice map in microservices_architecture.png.
See the Architecture reference for the component inventory, the roles that deploy them, and how this repo plugs into the platform.
Deploys the RAG stack on a single node with defaults. In three steps you will have an assistant answering questions about documents you ingest.
Note
Prerequisites: a Xeon host with 60 logical cores, 128 GB RAM, and 200 GB free disk; Ubuntu 22.04/24.04; passwordless sudo; internet access; and Hugging Face access to the default models. Full list, including the 32-core limited deployment → Prerequisites.
init erag clones this repo and the inference layer it depends on, at the revisions the platform pins, and seeds the configuration for your chosen pipeline. install erag then pulls in the layers below it (infrastructure, platform, inference) and deploys the RAG application on top.
git clone https://github.com/intel/enterprise-ai-solutions.git
cd enterprise-ai-solutions
./es_auto_installer.sh configure # one-time machine prep (Python 3.11+, yq, kubectl, helm)
./es_auto_installer.sh init erag # clone + seed the erag layer and its dependencies
./es_auto_installer.sh install erag # deploy everythingSettings for this layer land in env/local/config.erag.yaml. Pick a different pipeline with --flavour:
./es_auto_installer.sh init erag --flavour docsumNote
For any parameters, customization options or multinode deployment, refer to documentation.
Tip
--env defaults to local. Tear down with ./es_auto_installer.sh teardown erag.
Install and teardown are environment-scoped: if you installed with --env prod, you must
tear down with --env prod.
The gateway binds ports 80 and 443 on the node, so no port forwarding is needed. Add each subdomain to /etc/hosts on the machine you browse from - wildcards do not work there:
<node-ip> solutions.ai grafana.solutions.ai keycloak.solutions.ai s3.solutions.ai seaweedfs.solutions.ai
Then open https://solutions.ai. First-login credentials are written to env/local/logs/rag/default_credentials.txt; you will be asked to change the password immediately.
Important
With the default self-signed certificates, visit https://s3.solutions.ai once and accept
the warning before ingesting documents. Not needed with custom certificates.
Sign in as the admin user, open the Admin Panel → Data Ingestion tab, and upload a file or point it at a URL. Once ingestion reports complete, ask a question in the chat and the answer will cite your document.
To verify the pipeline from the command line instead:
cd ext/enterprise.ai-erag/deployment
./scripts/test_connection.sh # ChatQnA; use test_docsum.sh or test_translation.sh for those flavours| Area | Component | What it gives you |
|---|---|---|
| Ingestion | Enhanced Data Preparation (EDP) | Extract, split, and embed documents from the object store or SharePoint, with opt-in scheduled sync |
| Retrieval | Vector database + reranking | Redis Cluster, PGVector, or Microsoft SQL Server backends, with opt-in per-user access control on the index |
| Pipelines | GMC operator + composer | ChatQnA, DocSum, AudioQnA, and translation flows, composed from shared steps |
| Safety | Guardrail microservices | Query filtering on by default; response filtering via the output_guard variant, ingestion filtering via edp_dp_guard_enabled |
| Access | Keycloak OIDC + APISIX | Single sign-on, realm roles per persona, and opt-in per-user document scoping |
| Agents | MCP gateway | Expose retrieval and ingestion to AI agents over Model Context Protocol |
Curious first? The demo below shows ChatQnA in action.
Note
The video showcases an earlier release. The current UI, installation flow, and feature set have moved on since it was recorded.
The Quick Start deploys the ChatQnA flavour with defaults. From here you can change the pipeline and its variants, swap models, enable multilingual retrieval, connect SharePoint or an external S3 store, wire up agents, and tune every microservice.
| Goal | Guide |
|---|---|
| Deploy step by step, endpoints and credentials | Deploy the RAG layer |
| Pipelines, flavours, and variants | Pipelines |
| Every configuration option | Configuration |
| Change the LLM, embedding, or reranking model | Models |
| SSO, MFA, and Active Directory federation | Authentication |
| Ingest from SharePoint Online | SharePoint |
| Connect AI agents | MCP Integration |
| External S3 or NetApp ONTAP document store | Object Store |
| Components, roles, and request flow | Architecture |
| Dashboards and logs | Telemetry |
| Scaling and tuning | Performance |
| Something isn't working | Troubleshooting |
| VMware deployment | Deploy on VMware |
| What it is and why it exists | Meet Intel® AI for Enterprise RAG |
| Common questions | FAQ |
| Terminology | Glossary |
- How to Integrate SharePoint Online with a RAG system
- Lenovo Validated Design: AI POD Mini for Enterprise RAG Implementation
- Give Your RAG a Voice: Building an Audio Q&A Experience with Intel® AI for Enterprise RAG
- Accelerate AI Value Creation with Nutanix and Intel® AI for Enterprise RAG
- Converging Paradigms: Architecting a Hybrid and Open Platform for Unified HPC and AI Workloads
- Starting With the End in Mind: Intel and Nutanix's Blueprint for an Enterprise-Grade RAG Chatbot
- Securing Enterprise RAG Deployments
- Document Summarization: Transforming Enterprise Content with Intel® AI for Enterprise RAG
- Scaling Intel® AI for Enterprise RAG Performance: 64-Core vs 96-Core Intel® Xeon®
- Comprehensive Analysis: Intel® AI for Enterprise RAG Performance
- Monitoring and Debugging RAG Systems in Production
- NetApp AIPod Mini - Deployment Automation
- Multi-node deployments using Intel® AI for Enterprise RAG
- Rethinking AI Infrastructure: How NetApp and Intel Are Unlocking the Future with AIPod Mini
- Deploying Scalable Enterprise RAG on Kubernetes with Ansible Automation
Submit questions, feature requests, and bug reports on the GitHub Issues page.
Intel® AI for Enterprise RAG is licensed under the Apache License Version 2.0. Refer to the "LICENSE" file for the full license text and copyright notice.
This distribution includes third-party software governed by separate license terms. This third-party software, even if included with the distribution of the Intel software, may be governed by separate license terms, including without limitation, third-party license terms, other Intel software license terms, and open-source software license terms. These separate license terms govern your use of the third-party programs as set forth in the "THIRD-PARTY-PROGRAMS" file.
Please note: component(s) depend on software subject to non-open source licenses. If you use or redistribute this software, it is your sole responsibility to ensure compliance with such licenses.
The Security Policy outlines our guidelines and procedures for ensuring the highest level of security and trust for our users who consume Intel® AI for Enterprise RAG.
Intel is committed to respecting human rights and avoiding complicity in human rights abuses. See Intel's Global Human Rights Principles. Intel's products and software are intended only to be used in applications that do not cause or contribute to a violation of an internationally recognized human right.
You, not Intel, are responsible for determining model suitability for your use case. For information regarding model limitations, safety considerations, biases, or other information consult the model cards (if any) for models you use, typically found in the repository where the model is available for download. Contact the model provider with questions. Intel does not provide model cards for third party models.
If you want to contribute to the project, please refer to the guide in CONTRIBUTING.md file.
- Documentation Index
- GitHub Repository
- Intel® AI for Enterprise Solutions (platform)
- Intel® Enterprise for AI Inference (inference layer)
- Architecture
Intel, the Intel logo, OpenVINO, the OpenVINO logo, Pentium, and Xeon are trademarks of Intel Corporation or its subsidiaries. Other names and brands may be claimed as the property of others.
© Intel Corporation