Red Hat Developer Hub 2.1

Interacting with Red Hat Developer Lightspeed for Red Hat Developer Hub

Leverage Artificial Intelligence (AI)-driven expertise of the Red Hat Developer Lightspeed for Red Hat Developer Hub (Developer Lightspeed for RHDH) virtual assistant to help you use Red Hat Developer Hub (RHDH)

Red Hat Customer Content Services

Abstract

Red Hat Developer Lightspeed for Red Hat Developer Hub (Developer Lightspeed for RHDH) is an AI-powered virtual assistant for Red Hat Developer Hub (RHDH). You can interact with Developer Lightspeed for RHDH to explore RHDH capabilities in detail.

Preface

Red Hat Developer Lightspeed for Red Hat Developer Hub (Developer Lightspeed for RHDH) is an AI-powered virtual assistant for Red Hat Developer Hub (RHDH). You can interact with Developer Lightspeed for RHDH to explore RHDH capabilities in detail.

Chapter 1. Chat assistance with Developer Lightspeed for RHDH

Use Developer Lightspeed for RHDH to find product information, discover features, and resolve technical questions using natural language prompts directly within the RHDH console.

Lightspeed chatbot on the home page

Chapter 2. Red Hat Developer Lightspeed for Red Hat Developer Hub architecture for your AI backend deployment

Review the Developer Lightspeed for RHDH component architecture to plan your system layout and coordinate connections with your artificial intelligence (AI) backend deployment.

The architecture relies on the Lightspeed Core Service (LCORE) container, which operates as the primary intermediary layer to manage Developer Lightspeed for RHDH functionality and console user interactions. After you enable the plugin, the interface appears as the Intelligent Assistant button on all platforms that host RHDH.

2.1. AI reference and tool-calling capabilities through Lightspeed Core Service

Review the core components managed by the Lightspeed Core Service(LCORE) sidecar container to plan integrations with large language models (LLM) and tool runtime providers.

The LCORE container deploys as a sidecar to extend RHDH functionality. The container integrates and manages the following core architectural components:

  • Large language model (LLM) inference providers
  • Model Context Protocol (MCP) or Retrieval Augmented Generation (RAG) tool runtime providers

    Important

    Verify that your model supports tool calling before you enable MCP features. Using an incompatible model results in error messages.

  • Safety providers
  • Vector database settings

LCORE also manages critical operational configuration and key data, specifically:

  • User feedback collection
  • MCP server configuration
  • Chat history

Developer Lightspeed for RHDH sends prompts and receives LLM responses through the LCORE sidecar.

Chapter 3. Retrieval augmented generation (RAG) embeddings for grounded AI responses

Use retrieval-augmented generation (RAG) embeddings to ground artificial intelligence (AI) responses in your internal documentation and provide verified citations during user interactions.

The RHDH documentation serves as the primary data source for RAG operations. To provide accurate citations to production documentation during inference, the system uses RAG embeddings stored within a vector database.

The system processes RAG data through the following sequence:

  • An initialization container copies the RAG data to a shared volume.
  • The Lightspeed Core Service (LCORE) sidecar container mounts the shared volume to access the data.
  • The sidecar layer uses the embeddings to attach precise documentation references to the chat responses.

Chapter 4. Configure Developer Lightspeed for RHDH to initialize the AI assistant

Red Hat Developer Lightspeed for Red Hat Developer Hub is enabled by default on Red Hat Developer Hub (RHDH) instances. To provide developers with chat assistance, configure your deployment settings by using either the Operator or the Helm chart.

4.1. Configure the virtual assistant components

Configure Developer Lightspeed for RHDH by updating your Backstage custom resource (CR) to map environment variables, manage configurations, and set access rights.

Prerequisites

  • The RHDH Operator is installed on your cluster.
  • You have cluster administrator privileges.

Procedure

  1. Create an opaque Kubernetes Secret containing your operational credentials and query safety guardrails before applying the Backstage CR. Refer to the following key definitions for required environment variables:

    Important

    To disable an inference provider or configuration feature, you must leave the corresponding ENABLE_* variable completely unset. Setting an ENABLE_* variable to false does not disable the component because the underlying system checks only whether the variable is defined.

    KeyDescription

    ENABLE_VLLM

    Enables the vLLM platform when set to "true".

    VLLM_URL

    Specifies the target API endpoint URL for vLLM (for example, https://<api_endpoint>/v1).

    VLLM_API_KEY

    Stores the authorization token for your vLLM platform.

    ENABLE_OPENAI

    Enables the OpenAI platform when set to "true".

    OPENAI_API_KEY

    Stores the authorization secret key for OpenAI.

    ENABLE_VERTEX_AI

    Enables the Vertex AI platform when set to "true".

    VERTEX_AI_PROJECT

    Specifies your Google Cloud project ID.

    VERTEX_AI_LOCATION

    Specifies your target Google Cloud region.

    GOOGLE_APPLICATION_CREDENTIALS

    Specifies the file path of your mounted Google Cloud service account credentials JSON file.

    ENABLE_VALIDATION

    Activates query safety validation guardrails when set to "true".

    VALIDATION_PROVIDER

    Defines the active provider managing the verification routines (for example, openai or vllm).

    VALIDATION_MODEL_NAME

    Specifies the exact verification model to use (for example, gpt-4o-mini).

    The following code shows an example configuration Secret for vLLM with validation:

    apiVersion: v1
    kind: Secret
    metadata:
      name: lightspeed-auth-secrets
    type: Opaque
    stringData:
      ENABLE_VLLM: "true"
      VLLM_URL: "https://<api_endpoint>/v1"
      VLLM_API_KEY: "<api_key>"
      ENABLE_VALIDATION: "true"
      VALIDATION_PROVIDER: "vllm"
      VALIDATION_MODEL_NAME: "llama3.1"
  2. Map your secret inside the extraEnvs section of the Backstage CR to complete container provisioning:

    apiVersion: rhdh.redhat.com/v1alpha5
    kind: Backstage
    metadata:
      name: lightspeed-rhdh
    spec:
      application:
        extraEnvs:
          secrets:
            - name: lightspeed-auth-secrets
              containers:
                - lightspeed-core
  3. Optional: To protect settings such as Model Context Protocol (MCP) server additions from being overwritten during reconciliation loops, define a custom ConfigMap mapping in the extraFiles section of the CR:

        extraFiles:
          configMaps:
            - name: "my-custom-config"
              mountPath: /app-root
              key: lightspeed-stack.yaml
              containers:
                - lightspeed-core
  4. Configure access rights by updating the RBAC policy inside your Backstage CR:

    To grant non-administrator teams access to the virtual assistant, append permission lines to the rbac-policies.csv section, replacing <team> with your target team name:

    p, role:default/<team>, intelligent-assistant.chat, use, allow 1
    <1>

    Grants full use of the Developer Lightspeed for RHDH chat feature, including opening the chat window, sending messages, and managing chat history.

    For the complete list of AI feature permissions, including Notebooks, MCP tools, and skills, see AI feature permissions.

  5. Apply the updated custom resource manifest to your cluster:

    $ oc apply -f <backstage_cr_file>.yaml

Verification

  1. Log in to your console instance.
  2. Verify that the Intelligent Assistant button appears on the home page.
  3. Select the Intelligent Assistant button and confirm that the chat window initializes successfully.

4.2. Configure Developer Lightspeed for RHDH by using the Helm chart

Configure Developer Lightspeed for RHDH by using the Helm chart to manage large language model (LLM) providers, enable validation guardrails, and authorize custom role-based access control (RBAC) policies.

Prerequisites

  • You have access to a running RHDH instance deployed with Helm.
  • You have operational credentials for your chosen LLM provider.

Procedure

  1. Create a manual Kubernetes Secret to store your provider credentials by adding the required keys to your Secret based on your provider requirements:

    Important

    To disable an inference provider or configuration feature, you must leave the corresponding ENABLE_* variable completely unset. Setting an ENABLE_* variable to false does not disable the component because the underlying system checks only whether the variable is defined.

    KeyDescription

    ENABLE_VLLM

    Enables the vLLM platform when set to "true".

    VLLM_URL

    Specifies the target API endpoint URL for vLLM (for example, https://<api_endpoint>/v1).

    VLLM_API_KEY

    Stores the authorization token for your vLLM platform.

    ENABLE_OPENAI

    Enables the OpenAI platform when set to "true".

    OPENAI_API_KEY

    Stores the authorization secret key for OpenAI.

    ENABLE_VERTEX_AI

    Enables the Vertex AI platform when set to "true".

    VERTEX_AI_PROJECT

    Specifies your Google Cloud project ID.

    VERTEX_AI_LOCATION

    Specifies your target Google Cloud region.

    GOOGLE_APPLICATION_CREDENTIALS

    Specifies the file path of your mounted Google Cloud service account credentials JSON file.

    ENABLE_VALIDATION

    Activates query safety validation guardrails when set to "true".

    VALIDATION_PROVIDER

    Defines the active provider managing the verification routines (for example, openai or vllm).

    VALIDATION_MODEL_NAME

    Specifies the exact verification model to use (for example, gpt-4o-mini).

    Note
    • By default, the Helm installation creates a temporary Kubernetes Secret containing keys for various LLM providers. On subsequent helm upgrade cycles, the system overwrites this default Secret. Create a manual Kubernetes Secret to persist your credentials.
    • Vertex AI requires custom architecture mapping and has received limited testing.
    Tip

    To filter and reject off-topic user queries, you can optionally configure query safety validation guardrails within this Secret by defining the ENABLE_VALIDATION, VALIDATION_PROVIDER, and VALIDATION_MODEL_NAME keys.

  2. Reference your manual secret inside the values.yaml file:

    global:
      lightspeed:
        secret:
          create: false
          name: "my-custom-secret"
  3. Optional: Configuration files such as lightspeed-stack.yaml, config.yaml and rhdh-profile.py are managed by the Helm deployment and overwrites changes on Helm upgrade runs. To protect changes to a file, create a custom config map and reference it in the values.yaml file:

    Important

    Only modify the create and nameOverride fields. Keep the default mount paths and file configurations unchanged.

    global:
      lightspeed:
        configMaps:
          - name: stack
            create: false
            nameOverride: "my-custom-stack"
            mountPath: /app-root/lightspeed-stack.yaml
            subPath: lightspeed-stack.yaml
            sourceFile: lightspeed-stack.yaml
            optional: false
  4. Configure access rights by updating your RBAC definitions:

    To grant non-administrator teams access to the virtual assistant, append permission lines to the rbac-policies.csv section, replacing <team> with your target team name:

    p, role:default/<team>, intelligent-assistant.chat, use, allow 1
    <1>

    Grants full use of the Developer Lightspeed for RHDH chat feature, including opening the chat window, sending messages, and managing chat history.

    For the complete list of AI feature permissions, including Notebooks, MCP tools, and skills, see AI feature permissions.

  5. Run the helm upgrade command to apply your configurations to the cluster.

Verification

  1. Log in to your console instance.
  2. Verify that the Intelligent Assistant button appears on the home page.
  3. Select the Intelligent Assistant button and confirm that the chat window initializes successfully.

4.3. Disable Developer Lightspeed for RHDH by using the Operator

Disable the Developer Lightspeed for RHDH chat interface and stop associated container processes to remove the service from your Operator-backed deployment.

Prerequisites

  • You have access to the cluster where your instance is deployed.
  • You have cluster administrator privileges.

Procedure

  1. Open your Backstage custom resource (CR) YAML file.
  2. In the spec section, set the enabled flag to false for the lightspeed flavour. This disables the chat interface and prevents the Operator from injecting unconfigured sidecar containers:

    apiVersion: rhdh.redhat.com/v1alpha5
    kind: Backstage
    metadata:
      name: lightspeed-disabled
    spec:
      flavours:
        - name: lightspeed
          enabled: false
  3. Apply the updated custom resource manifest to your cluster:

    oc apply -f <backstage_cr_file>.yaml

Verification

  1. Log in to your console instance.
  2. Verify that the Intelligent Assistant button no longer appears on the home page.

4.4. Disable Developer Lightspeed for RHDH by using the Helm chart

Disable the Developer Lightspeed for RHDH chat interface and stop associated container processes to remove the service from your Helm-deployed environment.

Prerequisites

  • You have access to the cluster where your instance is deployed.
  • You have cluster administrator privileges.

Procedure

  1. Open your Helm values.yaml file.
  2. Update the global.lightspeed.enabled parameter to false to disable the chat interface:

    global:
      lightspeed:
        enabled: false
  3. Run the helm upgrade command to apply the configuration change to your cluster.

Verification

  1. Log in to your console instance.
  2. Verify that the Intelligent Assistant button no longer appears on the home page.

4.5. Mirror Developer Lightspeed for RHDH images for air-gapped environments

To provide chat assistance in a network environment without internet access, you must mirror the required Developer Lightspeed for RHDH container images and dynamic plugins to your local registry. This ensures your secure environment can pull the necessary components inside your network perimeter.

4.5.1. Mirror Lightspeed images for air-gapped environments

Mirror the required Developer Lightspeed for RHDH container images and plugins to your local registry to provide chat assistance in an air-gapped environment.

The prepare-restricted-environment.sh script does not automatically parse Developer Lightspeed for RHDH images from the bundle manifest, so mirror these images manually before running the script.

Prerequisites

  • You have a target mirror registry accessible to your disconnected cluster.
  • You authenticated to the Red Hat Container Registry and your target mirror registry.
  • You have configured image pull authentication for your mirror registry as described in Install Red Hat Developer Hub in an air-gapped environment with the Operator. The kubelet requires these credentials to pull all Red Hat Developer Hub container images, including the Developer Lightspeed for RHDH sidecar images.

Procedure

  1. Extract the deployment configurations from the official Operator bundle:

    BUNDLE_IMAGE="registry.redhat.io/rhdh/rhdh-operator-bundle:2.1"
    CONTAINER_ID=$(podman create "${BUNDLE_IMAGE}")
    podman cp $CONTAINER_ID:/manifests/rhdh-flavour-lightspeed-config_v1_configmap.yaml ./lightspeed-config.yaml
    podman rm $CONTAINER_ID
  2. Identify the initialization and sidecar container image tags from the extracted configuration file:

    LS_RAG_IMAGE=$(yq '.data["deployment.yaml"]' lightspeed-config.yaml | yq '.spec.template.spec.initContainers[] | select(.name == "init-rag-data") | .image')
    LS_CORE_IMAGE=$(yq '.data["deployment.yaml"]' lightspeed-config.yaml | yq '.spec.template.spec.containers[] | select(.name == "lightspeed-core") | .image')
  3. Mirror the images to your internal registry by running the skopeo copy command:

    skopeo copy docker://${LS_RAG_IMAGE} docker://<mirror_registry>/<ls_rag_repo>@<digest>
    skopeo copy docker://${LS_CORE_IMAGE} docker://<mirror_registry>/<ls_core_repo>@<digest>

4.5.2. Mirror Developer Lightspeed for RHDH images for Helm deployments on OpenShift Container Platform

Mirror the required Developer Lightspeed for RHDH container images to your local registry by using the oc-mirror plugin when deploying the Helm chart on an OpenShift Container Platform cluster.

Prerequisites

Procedure

  1. Identify the initialization and sidecar container images from the default values file of the chart:

    helm show values redhat-developer-hub --repo https://charts.openshift.io/ --version 2.1.0 > values.default.yaml
    
    LS_RAG_IMAGE=$(yq '.global.lightspeed.initContainer.image | .registry + "/" + .repository' values.default.yaml)
    LS_RAG_DIGEST=$(yq '.global.lightspeed.initContainer.image.tag' values.default.yaml)
    LS_CORE_IMAGE=$(yq '.global.lightspeed.sidecar.image | .registry + "/" + .repository' values.default.yaml)
    LS_CORE_DIGEST=$(yq '.global.lightspeed.sidecar.image.tag' values.default.yaml)
  2. Add these images to the additionalImages section of your ImageSetConfiguration file:

    apiVersion: mirror.openshift.io/v2alpha1
    kind: ImageSetConfiguration
    mirror:
      additionalImages:
        - name: ${LS_RAG_IMAGE}:${LS_RAG_DIGEST}
        - name: ${LS_CORE_IMAGE}:${LS_CORE_DIGEST}
      helm:
        repositories:
          - name: openshift-charts
            url: https://charts.openshift.io
            charts:
              - name: redhat-developer-hub
                version: "2.1"

4.5.3. Mirror Developer Lightspeed for RHDH images for Helm deployments on Kubernetes

Isolate image references, copy them manually to your internal registry, and update your configuration file when deploying the Helm chart on non-OpenShift platforms.

Prerequisites

Procedure

  1. Extract the image references from the default values file of the chart:

    $ helm show values redhat-developer-hub --repo https://charts.openshift.io/ --version 2.1.0 > values.default.yaml
    
    LS_RAG_IMAGE=$(yq '.global.lightspeed.initContainer.image | .registry + "/" + .repository' values.default.yaml)
    LS_RAG_DIGEST=$(yq '.global.lightspeed.initContainer.image.tag' values.default.yaml)
    LS_CORE_IMAGE=$(yq '.global.lightspeed.sidecar.image | .registry + "/" + .repository' values.default.yaml)
    LS_CORE_DIGEST=$(yq '.global.lightspeed.sidecar.image.tag' values.default.yaml)
  2. Mirror the images to your internal mirror registry:

    skopeo copy --all docker://${LS_RAG_IMAGE}:${LS_RAG_DIGEST} docker://<mirror_registry_name>/<ls_rag_repo_name>:${LS_RAG_DIGEST}
    skopeo copy --all docker://${LS_CORE_IMAGE}:${LS_CORE_DIGEST} docker://<mirror_registry_name>/<ls_core_repo_name>:${LS_CORE_DIGEST}
  3. Update your custom Helm values file with the mirrored registry locations and plugin references. Mirror the dynamic plugins to the local registry before you add their package paths to the file. For mirroring instructions, see Mirroring Red Hat Developer Hub dynamic plugins in disconnected environments:

    global:
      lightspeed:
        initContainer:
          image:
            registry: "<mirror_registry_name>"
            repository: <ls_rag_repo_name>
            tag: "${LS_RAG_DIGEST}"
        sidecar:
          image:
            registry: "<mirror_registry_name>"
            repository: <ls_core_repo_name>
            tag: "${LS_CORE_DIGEST}"
        plugins:
          - package: "oci://<mirror_registry_name>/rhdh/red-hat-developer-hub-backstage-plugin-lightspeed@<ls_frontend_digest>"
            disabled: false
          - package: "oci://<mirror_registry_name>/rhdh/red-hat-developer-hub-backstage-plugin-lightspeed-backend@<ls_backend_digest>"
            disabled: false

Chapter 5. AI feature permissions

To control who can use each Developer Lightspeed for RHDH AI feature, assign the feature-linked RBAC permissions that grant access to the chat, Notebooks, MCP tools, and skills.

Each feature uses a single permission that grants full access of that feature. Permissions do not separate create, read, update, and delete actions: a role either has full permission to use a feature or no access. All AI feature permissions use the use action.

PermissionWhat it grants

intelligent-assistant.chat

Use the Developer Lightspeed for RHDH chat feature, including opening the chat window, sending messages, and managing chat history.

intelligent-assistant.notebooks

Use the Developer Lightspeed for RHDH Notebooks feature, including listing, reading, creating, uploading, querying, updating, and deleting Notebooks.

intelligent-assistant.mcp.tools

Use the Model Context Protocol (MCP) tools, including listing MCP servers and managing their configuration.

intelligent-assistant.skills

Use the Developer Lightspeed for RHDH skills feature.

Agent Skills permissions

PermissionActionWhat it grants

intelligent-assistant.skills

use

View information about the Agent Skills that are available to the deployment. This permission controls visibility only. It does not control whether skills are loaded or whether configured skills influence answers.

Example RBAC policy

The following example grants a team full access to all Developer Lightspeed for RHDH AI features:

p, role:default/<team>, intelligent-assistant.chat, use, allow
p, role:default/<team>, intelligent-assistant.notebooks, use, allow
p, role:default/<team>, intelligent-assistant.mcp.tools, use, allow
p, role:default/<team>, intelligent-assistant.skills, use, allow

To grant access to a subset of features, include permissions for only those features.

Migration from previous permission names

If you are upgrading from Developer Hub 1.x, replace the previous lightspeed.* permission names with the feature-linked names. Developer Hub 2.1 grants each feature with a single permission and no longer uses individual action verbs such as create, read, update, or delete.

Previous permissions (1.x)New permission (2.1)

lightspeed.chat.read, lightspeed.chat.create, lightspeed.chat.update, lightspeed.chat.delete

intelligent-assistant.chat

lightspeed.notebooks.use

intelligent-assistant.notebooks

lightspeed.mcp.read, lightspeed.mcp.manage

intelligent-assistant.mcp.tools

Important

This is a breaking change. Developer Hub 2.1 does not support the previous permission names. You must update all RBAC policies, conditional rules, and automation scripts that reference the previous lightspeed.* permission identifiers before you upgrade.

Chapter 6. Extend Developer Hub intelligent assistant with Bring Your Own Knowledge (BYOK)

Add your organization’s internal documentation as a retrieval-augmented generation (RAG) knowledge source so that Developer Hub intelligent assistant can search and cite your content alongside the default Red Hat Developer Hub (RHDH) product documentation.

6.1. Bring Your Own Knowledge (BYOK) for Developer Hub intelligent assistant

Bring Your Own Knowledge (BYOK) enables Developer Hub intelligent assistant to retrieve information from documentation that your organization provides. An administrator generates an OGX vector store from source documents and registers it as a retrieval source.

By default, Developer Hub intelligent assistant grounds AI responses in Red Hat Developer Hub (RHDH) product documentation. With BYOK, you can add additional knowledge sources — such as internal runbooks, architecture guides, or custom API references — so that Developer Hub intelligent assistant can retrieve and cite content that is specific to your organization.

Note

This workflow is intended for RHDH administrators. The upstream rag-content README is the source of truth for supported vector-store generation commands and current prerequisites. See the additional resources.

6.1.1. How BYOK works

The BYOK workflow has four parts:

  1. Prepare source documents and citation metadata.
  2. Generate embeddings and an OGX vector store with rag-content.
  3. Deliver the vector database and, when required, a local embedding model to the Developer Hub intelligent assistant runtime.
  4. Register the store as a retrieval source and verify it through the Developer Hub intelligent assistant.
Important

The embedding model and dimension used to generate the database must match the model and dimension in the runtime configuration. A mismatch prevents reliable retrieval and can stop the vector store from loading.

6.1.2. Retrieval source ordering and prioritization

The order of entries under rag.retrieval.inline.sources or rag.retrieval.tool.sources does not assign priority to the vector stores. For example, listing custom-docs before okp does not guarantee that content from custom-docs is returned first.

Developer Hub intelligent assistant supports per-store BYOK weighting only with Inline RAG. Set score_multiplier on each entry under rag.byok.stores to adjust its relative importance. Values greater than 1.0 boost that store’s results, while values less than 1.0 reduce them.

Inline RAG queries the configured BYOK stores, multiplies each raw relevance score by that store’s score_multiplier, merges the results, and ranks them by weighted score. Consequently, a store with a higher score_multiplier receives a boost regardless of where it appears in the sources list.

Tool RAG does not apply score_multiplier. It exposes the configured stores through the file_search tool. OGX searches them, merges their chunks, and ranks the combined results by the scores returned by the stores. Use Tool RAG when the model should decide when to search. Use Inline RAG when every request must retrieve context and you require configurable BYOK store weighting.

Note

score_multiplier applies only to BYOK stores. It does not weight OKP results, whose scores use a different scoring system. When you combine BYOK and OKP through Inline RAG, enable the supported reranker to normalize and rerank results across those sources.

6.2. Prepare the source documents for BYOK RAG

Organize the documents that you want Developer Hub intelligent assistant to search into a directory, and add citation metadata before you generate the vector store.

The examples in this workflow use Markdown documents, an OGX FAISS store, and the sentence-transformers/all-mpnet-base-v2 embedding model. This model produces embeddings with dimension 768.

Prerequisites

  • You have a workstation with Git and uv installed, or a container runtime such as Podman.
  • You have access to the lightspeed-core rag-content repository.
  • You have source documents that you are authorized to use.
  • You have access to a container registry or another approved method of delivering files to your RHDH deployment.
  • You have a Developer Hub intelligent assistant deployment configured with the sentence_transformers inference provider.
  • You have network access to download the embedding model from Hugging Face, unless the model is supplied locally.

Procedure

  1. Create a directory for the documents that you want the assistant to search. For example:

    custom_docs/
      installation-guide.md
      operations-guide.md

    Use clear headings and ordinary text wherever possible. Text embedded only in screenshots or scanned PDF pages must be converted with optical character recognition before it can be indexed reliably.

  2. Process one document type per generation run and set --doc-type accordingly. This example uses markdown.
  3. For Markdown content, add YAML frontmatter to identify the document and its canonical source URL:

    ---
    title: Example Operations Guide
    url: https://docs.example.com/operations-guide
    ---
    
    # Example Operations Guide
    
    Document content begins here.

    The metadata is stored with each generated chunk. Developer Hub intelligent assistant can use the title and URL when presenting retrieval sources and citations.

    Important

    Do not include secrets, credentials, personal data, or content that users of the assistant are not authorized to retrieve. Access controls on the source system are not automatically reproduced inside a vector store.

6.3. Generate a vector database for BYOK RAG

Install the rag-content tooling, generate a portable OGX FAISS vector store from your custom documentation, and test retrieval before you deploy the store to Developer Hub intelligent assistant.

Prerequisites

Procedure

  1. Clone the rag-content repository and install its pinned dependencies:

    $ git clone https://github.com/lightspeed-core/rag-content.git
    $ cd rag-content
    $ uv sync
    Tip

    Use a specific release or commit in repeatable production workflows. The generator, query helper, and runtime dependencies must be kept compatible.

  2. From the rag-content repository, generate a portable OGX FAISS vector store:

    $ uv run python scripts/generate_embeddings.py \
      --folder /path/to/custom_docs \
      --output /path/to/vector_db/custom_docs \
      --index custom-docs \
      --vector-store llamastack-faiss \
      --model-dir "" \
      --model-name sentence-transformers/all-mpnet-base-v2 \
      --doc-type markdown \
      --chunk-size 512 \
      --chunk-overlap 128

    The llamastack-faiss value is retained as a compatibility name in the rag-content command-line interface. Current rag-content releases use OGX to create this store.

    Passing an empty value to --model-dir records the Hugging Face model ID rather than a workstation-specific model path. This makes the generated store portable. At runtime, the same model ID is resolved by the sentence_transformers provider.

    Note

    This model-ID configuration can require network access when the model is not already cached. For a disconnected deployment, package the embedding model and configure its runtime filesystem path as described in Deploy BYOK RAG in a disconnected environment.

  3. Confirm the output files. The output directory contains:

    • faiss_store.db: the vector database required at runtime.
    • lightspeed-stack.yaml: an example Developer Hub intelligent assistant configuration generated for the store.
    • llama-stack.yaml: the OGX configuration used to create the store.
  4. Read the unique vector-store ID from the generated configuration:

    $ yq '.registered_resources.vector_stores[0].vector_store_id' \
      /path/to/vector_db/custom_docs/llama-stack.yaml

    Save this value. The runtime configuration must use the exact same ID.

  5. Test retrieval before deployment by using the query helper from the same rag-content checkout:

    $ uv run python scripts/query_rag.py \
      --db-path /path/to/vector_db/custom_docs \
      --product-index custom-docs \
      --model-path "" \
      --top-k 5 \
      --query "Enter a question answered only by the custom documents" \
      --json

    Review the returned chunks, scores, titles, and source URLs. Use a question that contains a distinctive fact from the custom content so that the result is easy to distinguish from general model knowledge.

    Note

    Some combinations of OGX and rag-content can report that the vector store is already registered when query_rag.py opens a newly generated database. Do not edit the SQLite database manually. Use a compatible rag-content revision or continue validation through the deployed Developer Hub intelligent assistant vector-store API, and report reproducible helper failures to the rag-content project.

6.4. Package a BYOK RAG container image

Bundle your generated vector store into a container image so that it can be delivered to your Red Hat Developer Hub (RHDH) cluster and copied into the Developer Hub intelligent assistant runtime by an init container.

The RHDH deployment can use an init container to copy the database from the image into a shared volume that the Developer Hub intelligent assistant container mounts. Persistent volumes, configuration-management systems, or other organization-approved delivery mechanisms can also be used.

Prerequisites

  • You have generated a portable OGX FAISS vector store as described in Generate a vector database for BYOK RAG.
  • You have podman installed.
  • You have write access to a container registry accessible to your cluster.

Procedure

  1. Create a Containerfile in the directory that contains your vector_db output. Use the variant that matches your deployment method:

    For the Operator, or any deployment that adds a dedicated BYOK init container alongside the product init container, use a minimal base image:

    FROM registry.access.redhat.com/ubi9-micro:latest
    COPY vector_db /byok/vector_db
    USER 1001

    For the Helm chart, which exposes a single global.lightspeed.initContainer that you must override, derive the image from the product RAG init image so that the overridden init container can deliver both the product-provided content and your BYOK vector store:

    ARG PRODUCT_RAG_IMAGE
    FROM ${PRODUCT_RAG_IMAGE}
    
    COPY vector_db /byok/vector_db
    USER 1001

    Set PRODUCT_RAG_IMAGE to the value of global.lightspeed.initContainer.image from your installed chart version. The derived image retains the product /rag directory and adds your custom database under /byok/vector_db.

  2. Build the image:

    For the minimal UBI-based image:

    $ podman build --platform linux/amd64 \
      -t quay.io/your-organization/rhdh-byok:1.0.0 .

    For the Helm image derived from the product RAG init image, pass an immutable product image reference:

    $ podman build --platform linux/amd64 \
      --build-arg PRODUCT_RAG_IMAGE=<product_rag_image_reference> \
      -t quay.io/your-organization/rhdh-byok-with-product-rag:1.0.0 .
  3. Push the image to an approved registry:

    $ podman push quay.io/your-organization/rhdh-byok:1.0.0
    Important

    Use an immutable version tag or image digest in production. Ensure that the image architecture matches the OpenShift worker architecture.

  4. Note the final path visible inside the Developer Hub intelligent assistant container. The db_path configuration must match that runtime path exactly. For example, if the database is copied to /rag-content/vector_db/custom_docs/faiss_store.db, use that complete path in the Developer Hub intelligent assistant configuration.

6.5. Configure BYOK RAG by using the Helm chart

Deploy your BYOK container image as an init container and register the custom knowledge source in lightspeed-stack.yaml in your Helm-based Red Hat Developer Hub (RHDH) deployment.

Prerequisites

Procedure

  1. Prepare a complete lightspeed-stack.yaml file with the BYOK sections added, then create a ConfigMap from it:

    Warning

    The mounted file replaces the product-provided lightspeed-stack.yaml in its entirety; it is not merged. If you create the ConfigMap from only the BYOK sections, you remove other supported or required settings, including the service, authentication, cache, user-data and profile, MCP, and other product configuration, and Developer Hub intelligent assistant can fail to start or lose required functionality. Always create the ConfigMap from the complete configuration for your RHDH version with the BYOK sections added to it.

    1. Copy or export the complete product-provided lightspeed-stack.yaml for your installed RHDH version to a local file.
    2. Ensure that the sentence_transformers embedding provider is present under inference.providers, then register the store under rag.byok.stores and activate it under rag.retrieval. Add the following fragment to the local file, merging it with the existing sections rather than replacing them:

      inference:
        providers:
          - type: sentence_transformers  # Add only if not already present
      rag:
        byok:
          stores:
            - rag_id: custom-docs  1
              backend: faiss
              embedding_model: sentence-transformers/all-mpnet-base-v2
              embedding_dimension: 768
              vector_db_id: vs_replace_with_generated_id  2
              db_path: /rag-content/vector_db/custom_docs/faiss_store.db  3
              score_multiplier: 1.0  # Used by Inline RAG only
        retrieval:
          tool:
            sources:
              - custom-docs  4
      <1>

      The rag_id identifies the retrieval source. Use the same value under rag.retrieval.tool.sources.

      1 1 1
      Replace with the vector_db_id value that you read from the generated llama-stack.yaml file. The runtime configuration must use the exact same ID.
      2
      Must match the path where the init container copies the database inside the Developer Hub intelligent assistant container.
      3
      Must match rag_id. If you have other sources, such as product documentation, list each required source under sources.
    3. Create the ConfigMap from the completed local file:

      $ kubectl create configmap lightspeed-stack-byok \
          --from-file=lightspeed-stack.yaml \
          --dry-run=client -o yaml | kubectl apply -f -
  2. Update your Helm values.yaml file to reference the custom ConfigMap and override the Developer Hub intelligent assistant init container so that it delivers both the product-provided RAG content and your BYOK vector store. The chart mounts the RAG volume (global.lightspeed.ragVolume) automatically into both the init container and the Developer Hub intelligent assistant container at /rag-content, so you do not define extra volumes or volume mounts:

    Important

    The global.lightspeed.configMaps value is a Helm list. If you provide only the customized stack entry, you replace the entire default list and remove the config and rhdh-profile entries. Retain every entry from your installed chart version and change only the stack entry to reference lightspeed-stack-byok. The following example shows the minimum entries to retain; match the exact list to your chart version.

    global:
      lightspeed:
        configMaps:
          - name: stack
            create: false
            nameOverride: lightspeed-stack-byok  1
            mountPath: /app-root/lightspeed-stack.yaml
            subPath: lightspeed-stack.yaml
            sourceFile: lightspeed-stack.yaml
            optional: false
          - name: config
            create: true
            nameOverride: ""
            mountPath: /app-root/config.yaml
            subPath: config.yaml
            sourceFile: config.yaml
            optional: false
          - name: rhdh-profile
            create: true
            nameOverride: ""
            mountPath: /app-root/rhdh-profile.py
            subPath: rhdh-profile.py
            sourceFile: rhdh-profile.py
            optional: false
        initContainer:  2
          image: quay.io/your-organization/rhdh-byok-with-product-rag:1.0.0
          command: ["sh", "-c"]
          args:
            - |
              set -eu
              mkdir -p /tmp/data /rag-content/vector_db
              cp -R --no-preserve=mode,ownership \
                /rag/vector_db/. /rag-content/vector_db/
              cp -R --no-preserve=mode,ownership \
                /rag/embeddings_model /rag-content/
              cp -R --no-preserve=mode,ownership \
                /byok/vector_db/. /rag-content/vector_db/
              mkdir -p /rag-content/vector_db/notebooks
              chmod -R a+rwX \
                /rag-content/embeddings_model /rag-content/vector_db
    Important

    The chart exposes a single global.lightspeed.initContainer. Overriding it replaces the default init container that delivers the product-provided RAG content. The combined image built from the product RAG init image is therefore required so that the overridden init container can deliver both the product-provided content and your BYOK vector store. Give each custom store its own directory under vector_db so that it does not overwrite a product store. Validate the result against your chart version before relying on it in production.

    <1>

    Replace the default lightspeed-stack.yaml ConfigMap with your custom one that contains the BYOK configuration.

    1
    Override the default init container so that it copies both the product-provided content and your BYOK vector database into the automatically mounted /rag-content RAG volume before Developer Hub intelligent assistant starts.
  3. Run helm upgrade to apply the changes:

    $ helm upgrade <release-name> redhat-developer-hub \
        --repo https://charts.openshift.io/ \
        --version 2.1.0 \
        -f values.yaml

6.6. Configure BYOK RAG by using the Operator

Deploy your BYOK container image as an init container and register the custom knowledge source in lightspeed-stack.yaml in your Operator-based Red Hat Developer Hub (RHDH) deployment.

Prerequisites

Procedure

  1. Prepare a complete lightspeed-stack.yaml file with the BYOK sections added, then create a ConfigMap from it:

    Warning

    The mounted file replaces the product-provided lightspeed-stack.yaml in its entirety; it is not merged. If you create the ConfigMap from only the BYOK sections, you remove other supported or required settings, including the service, authentication, cache, user-data and profile, MCP, and other product configuration, and Developer Hub intelligent assistant can fail to start or lose required functionality. Always create the ConfigMap from the complete configuration for your RHDH version with the BYOK sections added to it.

    1. Copy or export the complete product-provided lightspeed-stack.yaml for your installed RHDH version to a local file.
    2. Ensure that the sentence_transformers embedding provider is present under inference.providers, then register the store under rag.byok.stores and activate it under rag.retrieval. Add the following fragment to the local file, merging it with the existing sections rather than replacing them:

      inference:
        providers:
          - type: sentence_transformers  # Add only if not already present
      rag:
        byok:
          stores:
            - rag_id: custom-docs  1
              backend: faiss
              embedding_model: sentence-transformers/all-mpnet-base-v2
              embedding_dimension: 768
              vector_db_id: vs_replace_with_generated_id  2
              db_path: /rag-content-byok/vector_db/custom_docs/faiss_store.db  3
              score_multiplier: 1.0  # Used by Inline RAG only
        retrieval:
          tool:
            sources:
              - custom-docs  4
      <1>

      The rag_id identifies the retrieval source. Use the same value under rag.retrieval.tool.sources.

      1
      Replace with the vector_db_id value that you read from the generated llama-stack.yaml file. The runtime configuration must use the exact same ID.
      2
      Must match the path where the init container copies the database inside the Developer Hub intelligent assistant container. This procedure uses the dedicated /rag-content-byok mount path so that the BYOK content does not replace or shadow any existing product-provided RAG content in /rag-content.
      3
      Must match rag_id. If you have other sources, such as product documentation, list each required source under sources.
    3. Create the ConfigMap from the completed local file:

      $ oc create configmap lightspeed-stack-byok \
          --from-file=lightspeed-stack.yaml \
          --dry-run=client -o yaml | oc apply -f -
  2. Update your Backstage custom resource (CR) to mount the custom ConfigMap, and use spec.deployment.patch to add the BYOK init container, shared volume, and volume mount. The spec.deployment.patch value is a deployment fragment that the Operator merges with the generated deployment, so the BYOK init container is added alongside any existing init containers, and the dedicated /rag-content-byok mount does not replace or shadow any existing product-provided RAG content:

    apiVersion: rhdh.redhat.com/v1alpha5
    kind: Backstage
    metadata:
      name: lightspeed-rhdh
    spec:
      application:
        extraFiles:
          configMaps:
            - name: "lightspeed-stack-byok"  1
              mountPath: /app-root
              key: lightspeed-stack.yaml
              containers:
                - lightspeed-core
      deployment:
        patch:
          spec:
            template:
              spec:
                volumes:
                  - name: byok-rag
                    emptyDir: {}
                initContainers:
                  - name: byok-rag-init  2
                    image: quay.io/your-organization/rhdh-byok:1.0.0
                    imagePullPolicy: IfNotPresent
                    command: ["/bin/sh", "-c"]
                    args:
                      - |
                        set -eu
                        mkdir -p /rag-content-byok/vector_db
                        cp -R /byok/vector_db/. /rag-content-byok/vector_db/
                    volumeMounts:
                      - name: byok-rag
                        mountPath: /rag-content-byok
                    securityContext:
                      allowPrivilegeEscalation: false
                      readOnlyRootFilesystem: true
                      runAsNonRoot: true
                      capabilities:
                        drop: ["ALL"]
                      seccompProfile:
                        type: RuntimeDefault
                containers:
                  - name: lightspeed-core
                    volumeMounts:
                      - name: byok-rag
                        mountPath: /rag-content-byok
    <1>

    References the ConfigMap created in the previous step. This replaces the default lightspeed-stack.yaml for the lightspeed-core container.

    1
    The init container runs before Developer Hub intelligent assistant starts and copies the vector database to the shared byok-rag volume. The dedicated /rag-content-byok mount path does not replace or shadow any existing product-provided RAG content.
  3. Apply the updated Backstage CR:

    $ oc apply -f <backstage_cr_file>.yaml

6.7. Deploy BYOK RAG in a disconnected environment

Package the embedding model alongside the vector store and register the store with local runtime paths so that Developer Hub intelligent assistant can retrieve custom content without network access to Hugging Face.

A precomputed FAISS database contains the document embeddings, but Developer Hub intelligent assistant must still embed each incoming question before it can search that database. The runtime therefore requires the same embedding model and dimension that were used to generate the store. Supplying only faiss_store.db is not sufficient for a disconnected deployment.

The rag-content project provides a model-download helper and also bundles its default embedding model in its RAG tool image. The following upstream references are pinned to a specific revision so that the examples do not change unexpectedly:

Prerequisites

  • You have generated a portable OGX FAISS vector store as described in Generate a vector database for BYOK RAG.
  • You have a connected build workstation that uses the same pinned rag-content checkout used to generate the vector store.
  • You have podman installed and a registry that is available to the disconnected OpenShift cluster.

Procedure

  1. On a connected build workstation, download the embedding model:

    $ mkdir -p embeddings_model
    
    $ uv run python scripts/download_embeddings_model.py \
      --local-dir ./embeddings_model \
      --hf-repo-id sentence-transformers/all-mpnet-base-v2

    The resulting directory includes the model weights, tokenizer, configuration, pooling, and normalization files needed by the sentence_transformers provider. Scan and transfer this directory according to your organization’s disconnected-content process.

    Note

    If the vector store is generated in a disconnected build environment, pass the local directory to --model-dir instead of an empty value. In either case, use the exact same model weights for generation and runtime. Changing the model requires regenerating the vector store.

  2. Place the generated vector_db directory and downloaded embeddings_model directory in the container build context:

    byok-image/
      Containerfile
      vector_db/
        custom_docs/
          faiss_store.db
      embeddings_model/
        model.safetensors
        config.json
        tokenizer.json
        ...
  3. Create an init-container image that contains both artifacts:

    FROM registry.access.redhat.com/ubi9/ubi-minimal:latest
    
    COPY vector_db /byok/vector_db
    COPY embeddings_model /byok/embeddings_model
    
    RUN chgrp -R 0 /byok && chmod -R g=u /byok
    USER 1001
    Important

    Pin the base image and resulting BYOK image by digest in production. The image reference is part of the Kubernetes deployment configuration. It is not an alternative value for db_path or embedding_model in the Developer Hub intelligent assistant BYOK configuration; those fields must identify paths visible inside the Developer Hub intelligent assistant container.

  4. Build the image on a connected system, scan it, and mirror it into a registry available to the disconnected OpenShift cluster:

    $ podman build -t registry.example.com/rhdh/byok-content:1.0.0 .
    $ podman push registry.example.com/rhdh/byok-content:1.0.0
  5. Deliver the artifacts to the Developer Hub intelligent assistant runtime and configure offline operation. The supported customization paths differ for the Helm chart and the Operator, so use the variant that matches your RHDH installation.

    Helm chart

    Because the chart exposes a single global.lightspeed.initContainer, derive a combined image from the product RAG init image that also includes your custom vector database and local embedding model. Use a Containerfile such as:

    ARG PRODUCT_RAG_IMAGE
    FROM ${PRODUCT_RAG_IMAGE}
    
    COPY vector_db /byok/vector_db
    COPY embeddings_model /byok/embeddings_model
    USER 1001

    Override the single global.lightspeed.initContainer so that it copies the product content, your BYOK database, and the local model into the automatically mounted /rag-content volume, and set the offline environment variables under global.lightspeed.sidecar.env:

    global:
      lightspeed:
        sidecar:
          env:
            - name: HF_HUB_OFFLINE
              value: "1"
            - name: TRANSFORMERS_OFFLINE
              value: "1"
        initContainer:
          image: quay.io/your-organization/rhdh-byok-with-product-rag:1.0.0
          command: ["sh", "-c"]
          args:
            - |
              set -eu
              mkdir -p /rag-content/vector_db
              cp -R --no-preserve=mode,ownership \
                /rag/vector_db/. /rag-content/vector_db/
              cp -R --no-preserve=mode,ownership \
                /byok/vector_db/. /rag-content/vector_db/
              cp -R --no-preserve=mode,ownership \
                /byok/embeddings_model /rag-content/embeddings_model
              chmod -R a+rwX \
                /rag-content/embeddings_model /rag-content/vector_db

    Operator

    Use spec.deployment.patch to add the byok-rag volume and the BYOK init container, mount the volume at /rag-content-byok in both the init container and lightspeed-core, copy both artifacts into it, and add the offline environment variables to lightspeed-core:

    spec:
      deployment:
        patch:
          spec:
            template:
              spec:
                volumes:
                  - name: byok-rag
                    emptyDir: {}
                initContainers:
                  - name: byok-rag-init
                    image: registry.example.com/rhdh/byok-content@sha256:replace_with_digest
                    imagePullPolicy: IfNotPresent
                    command: ["/bin/sh", "-c"]
                    args:
                      - |
                        set -eu
                        mkdir -p /rag-content-byok/vector_db /rag-content-byok/embeddings_model
                        cp -R /byok/vector_db/. /rag-content-byok/vector_db/
                        cp -R /byok/embeddings_model/. /rag-content-byok/embeddings_model/
                    volumeMounts:
                      - name: byok-rag
                        mountPath: /rag-content-byok
                    securityContext:
                      allowPrivilegeEscalation: false
                      readOnlyRootFilesystem: true
                      runAsNonRoot: true
                      capabilities:
                        drop: ["ALL"]
                      seccompProfile:
                        type: RuntimeDefault
                containers:
                  - name: lightspeed-core
                    env:
                      - name: HF_HUB_OFFLINE
                        value: "1"
                      - name: TRANSFORMERS_OFFLINE
                        value: "1"
                    volumeMounts:
                      - name: byok-rag
                        mountPath: /rag-content-byok

    Kubernetes init containers run before application containers, so the model and database are present when Developer Hub intelligent assistant starts.

    Note

    If several init containers populate the same volume, give each vector store a unique directory and ensure that later containers merge content instead of replacing the entire vector_db directory.

  6. Register the disconnected store with local runtime paths that match your deployment method:

    For the Helm chart:

    rag:
      byok:
        stores:
          - rag_id: custom-docs
            backend: faiss
            embedding_model: /rag-content/embeddings_model  1
            embedding_dimension: 768
            vector_db_id: vs_replace_with_generated_id
            db_path: /rag-content/vector_db/custom_docs/faiss_store.db
            score_multiplier: 1.0
      retrieval:
        tool:
          sources:
            - custom-docs

    For the Operator:

    rag:
      byok:
        stores:
          - rag_id: custom-docs
            backend: faiss
            embedding_model: /rag-content-byok/embeddings_model  1
            embedding_dimension: 768
            vector_db_id: vs_replace_with_generated_id
            db_path: /rag-content-byok/vector_db/custom_docs/faiss_store.db
            score_multiplier: 1.0
      retrieval:
        tool:
          sources:
            - custom-docs
    <1>
    Using the local path prevents the sentence_transformers provider from attempting to resolve sentence-transformers/all-mpnet-base-v2 from Hugging Face when the pod starts or processes its first query.

6.8. Verify the deployed vector store

Confirm that Developer Hub intelligent assistant can read the delivered artifacts, that the BYOK source is registered, and that the assistant retrieves and cites content from your custom documentation.

Prerequisites

Procedure

  1. Verify the following conditions after deployment:

    1. The init container or file-delivery process completes successfully.
    2. The Developer Hub intelligent assistant container can read the configured faiss_store.db path.
    3. For a disconnected deployment that uses a local embedding model, the Developer Hub intelligent assistant container can read the local embedding model files, including model.safetensors, the tokenizer, and model configuration. A connected deployment that uses a Hugging Face model ID does not require these local files.
    4. The Developer Hub intelligent assistant readiness endpoint reports success.
    5. The BYOK source appears in the RAG and vector-store API responses.
    6. The Developer Hub intelligent assistant answers a distinctive question from the custom documents and displays the expected source metadata.
  2. From inside the Developer Hub intelligent assistant container, test the exact db_path that you configured, then run the readiness and API checks. Use the path that matches your deployment method:

    # Helm example
    $ test -r /rag-content/vector_db/custom_docs/faiss_store.db
    
    # Operator example
    $ test -r /rag-content-byok/vector_db/custom_docs/faiss_store.db
    
    $ curl --fail http://127.0.0.1:8080/readiness
    $ curl --fail http://127.0.0.1:8080/v1/rags
    $ curl --fail http://127.0.0.1:8080/v1/vector-stores

    The exact host and port can differ depending on the installation. Protect administrative and diagnostic endpoints according to your organization’s security requirements.

    Note

    Run the following additional check only for a disconnected deployment that bundles a local embedding model. A connected deployment that uses a Hugging Face model ID, such as sentence-transformers/all-mpnet-base-v2, does not require this file. Test the configured embedding model path, which differs by deployment method:

    # Test <configured_embedding_model_path>/model.safetensors, for example:
    # Helm example
    $ test -r /rag-content/embeddings_model/model.safetensors
    
    # Operator example
    $ test -r /rag-content-byok/embeddings_model/model.safetensors
  3. In the Developer Hub intelligent assistant, ask a question whose answer appears only in the custom content. Confirm both the answer and the cited title or URL.

    Note

    A plausible answer without the expected retrieval source does not prove that BYOK retrieval occurred.

6.9. Update the BYOK RAG knowledge base

Regenerate and redeploy the vector store when your source documents change. The FAISS database is a generated artifact, so you must rebuild it from the authoritative document set rather than editing it in place.

Prerequisites

Procedure

  1. Regenerate the complete store from the authoritative document set.
  2. Record the newly generated vector-store ID.
  3. Build and publish a new immutable image or artifact version.
  4. Update the runtime vector_db_id, db_path when necessary, and image reference together.
  5. Roll out Developer Hub intelligent assistant and repeat the retrieval verification.

    Important

    Changing the embedding model requires regenerating the vector store. Do not reuse a database generated with a different model or embedding dimension.

6.10. BYOK RAG configuration fields for lightspeed-stack.yaml

Reference the inference, rag.byok, and rag.retrieval fields available in lightspeed-stack.yaml to register and activate custom knowledge sources for Developer Hub intelligent assistant.

6.10.1. inference fields

The sentence_transformers inference provider embeds each incoming question before Developer Hub intelligent assistant searches the vector store.

inference:
  providers:
    - type: sentence_transformers

6.10.2. rag.byok.stores list fields

Each entry in the rag.byok.stores list defines one custom knowledge source. You can define multiple sources.

FieldTypeRequiredDescription

rag_id

String

Yes

Unique identifier for this knowledge source. Referenced under rag.retrieval to activate the source.

backend

String

Yes

Vector store backend type. Use faiss for the OGX FAISS store generated by rag-content.

embedding_model

String

Yes

Hugging Face model ID for the embedding model, or a local path inside the Developer Hub intelligent assistant container for a disconnected deployment. Must match the model used during vector store generation. Default: sentence-transformers/all-mpnet-base-v2.

embedding_dimension

Integer

Yes

Output dimension of the embedding model. For sentence-transformers/all-mpnet-base-v2, use 768.

vector_db_id

String

Yes

Unique store identifier generated during vector store creation. Read this value from the generated llama-stack.yaml file. The runtime configuration must use the exact same ID.

db_path

String

Yes

Absolute path to the faiss_store.db file inside the Developer Hub intelligent assistant container. Must match the runtime path where the database is delivered.

score_multiplier

Float

No

Relative weight applied to this store’s results with Inline RAG. Values greater than 1.0 boost the store’s results; values less than 1.0 reduce them. Ignored by Tool RAG. Default: 1.0.

6.10.3. rag.retrieval fields

The rag.retrieval section activates configured BYOK sources and sets the retrieval mode.

FieldDescription

rag.retrieval.tool.sources

List of rag_id values exposed to the model through the file_search tool. The model retrieves from these sources when it determines retrieval is relevant. score_multiplier is not applied.

rag.retrieval.inline.sources

List of rag_id values whose results are retrieved for every request, weighted by score_multiplier, merged, and ranked by weighted score.

6.10.4. Example: Single source

inference:
  providers:
    - type: sentence_transformers

rag:
  byok:
    stores:
      - rag_id: custom-docs
        backend: faiss
        embedding_model: sentence-transformers/all-mpnet-base-v2
        embedding_dimension: 768
        vector_db_id: vs_replace_with_generated_id
        db_path: /rag-content/vector_db/custom_docs/faiss_store.db
        score_multiplier: 1.0  # Used by Inline RAG only
  retrieval:
    tool:
      sources:
        - custom-docs
Note

In this example, /rag-content/vector_db/…​ is the Helm chart path. The Operator procedure uses /rag-content-byok/vector_db/…​. The db_path value must always match the actual path mounted in the lightspeed-core container for your deployment method.

6.10.5. Example: Inline RAG with per-store weighting

rag:
  byok:
    stores:
      - rag_id: general-docs
        # Other required store settings are omitted for brevity.
        score_multiplier: 1.0
      - rag_id: preferred-docs
        # Other required store settings are omitted for brevity.
        score_multiplier: 1.2  1
  retrieval:
    inline:
      sources:
        - general-docs
        - preferred-docs
<1>
With Inline RAG, preferred-docs receives a boost regardless of where it appears in the sources list.

6.11. Troubleshooting BYOK RAG

Resolve common errors that occur when you configure or use a Bring Your Own Knowledge (BYOK) RAG knowledge source with Developer Hub intelligent assistant.

6.11.1. The store does not load

Symptom: Developer Hub intelligent assistant fails to load the vector store at startup or on the first query.

Resolution:

  • Confirm that db_path is the path inside the Developer Hub intelligent assistant container, not the path used on the generation workstation or inside the delivery image.
  • In a disconnected environment, confirm that embedding_model is the local path inside the Developer Hub intelligent assistant container rather than only a Hugging Face model ID.
  • Confirm that the init container copied the complete model directory, including weights, tokenizer, model configuration, pooling, and normalization files.
  • Confirm that vector_db_id exactly matches the generated ID.
  • Confirm that the database file is present, non-empty, and readable by the container user.
  • Review init-container and Developer Hub intelligent assistant startup logs.

6.11.2. Queries return irrelevant chunks

Symptom: The knowledge source is configured and citations appear, but retrieved content is empty or unrelated.

Resolution:

  • Confirm that the generation and runtime embedding models are identical.
  • Confirm that the configured embedding dimension matches the model.
  • Improve document headings and remove navigation or boilerplate text that dominates the content.
  • Adjust chunk size and overlap, regenerate the store, and compare results with representative test questions.

6.11.3. Citations are missing or incorrect

Symptom: Responses do not include the expected source title or URL, or cite the wrong source.

Resolution:

  • Confirm that each source document contains valid title and url frontmatter.
  • Inspect query results to ensure that title and URL metadata were stored with the chunks.
  • Confirm that the cited URL is accessible to the intended RHDH users.

Chapter 7. Customize Developer Lightspeed for RHDH AI responses

You can customize Developer Lightspeed for RHDH to align model behavior with your operational goals, enhance developer productivity, and ensure secure data retention.

Customize Developer Lightspeed for RHDH by enabling user feedback, persisting chat history, and configuring Model Context Protocol (MCP) tools.

7.1. Enable user feedback to improve model performance

Enable user feedback collection to allow users to rate chat responses and submit text comments directly within the console interface.

The Lightspeed Core Service (LCORE) stores this data as JSON files inside your cluster. Because Red Hat does not collect or access this data, platform administrators must manage, analyze, and delete these files locally.

Prerequisites

  • You have platform administrator privileges.
  • You created and referenced a custom config map in your deployment to ensure configuration changes persist during system upgrades or Operator reconciliation loops. For more information, see Provision your custom Red Hat Developer Hub configuration.

Procedure

  1. In your custom configuration file, such as lightspeed-stack.yaml, modify the user_data_collection block to configure your data preferences:

    • To enable feedback collection, set the feedback_enabled parameter to true:

      user_data_collection:
        feedback_enabled: true
        feedback_storage: "/tmp/data/feedback"
        transcripts_enabled: true
        transcripts_storage: "/tmp/data/transcripts"
    • To disable feedback collection, set the feedback_enabled parameter to false:

      user_data_collection:
        feedback_enabled: false
        feedback_storage: "/tmp/data/feedback"
        transcripts_enabled: true
        transcripts_storage: "/tmp/data/transcripts"
      Note

      Do not modify the feedback_storage or transcripts_storage data paths when disabling feedback. Altering these path strings prevents the service from locating existing historical logs.

  2. Apply the updated configuration file changes to your cluster by running your platform’s standard deployment or upgrade sequence.

7.2. Customize AI responses by using system prompts

Configure a custom system prompt to provide environmental context to the large language model (LLM). This custom instruction prefixes user queries, guiding the assistant to generate artificial intelligence (AI) responses tailored to your RHDH instance.

Prerequisites

  • You have administrative access to the RHDH host platform filesystem.

Procedure

  1. In your app-config.yaml file, add or modify the systemPrompt parameter under the lightspeed section, specifying your custom instruction string:

    lightspeed:
      # ... other lightspeed configurations
      systemPrompt: "You are a helpful assistant focused on Red Hat Developer Hub development."
  2. Save the file.
  3. Restart the RHDH service to apply the updated system prompt configuration.

7.3. Customize chat history storage

Configure chat history storage to choose between non-persistent local logs and a persistent external database for user conversations.

By default, the system stores chat history in a non-persistent local database within the Lightspeed Core Service (LCORE) container. To retain data across system restarts, you must configure a PostgreSQL database connection.

Warning

Storing chat history records user prompts and responses. You must assess data privacy and security implications if your user chat history contains private, sensitive, or confidential information. For users that want to have their chat data removed, they must request their platform administrator to perform this action. Red Hat does not collect or access this chat history data.

Prerequisites

Procedure

  1. In your custom configuration file, such as lightspeed-stack.yaml, modify the conversation_cache block to specify your storage configuration:

    • To enable persistent storage, add your PostgreSQL database credentials and endpoint properties:

      conversation_cache:
        type: "postgres"
        postgres:
          host: _<your_database_host>_
          port: _<your_database_port>_
          db: _<your_database_name>_
          user: _<your_user_name>"_
          password: _<postgres_password>_
      • To retain the default non-persistent SQLite setup, verify that the parameters match the following paths:

        conversation_cache:
          type: "sqlite"
          sqlite:
            db_path: '/tmp/cache.db'
  2. Restart the LCORE service to apply your new database configuration.

7.4. Configure rate limits for the Developer Hub intelligent assistant

Configure per-user rate limits for the Developer Hub intelligent assistant to prevent abuse, control large language model (LLM) inference costs, and protect backend resources.

Rate limiting is active by default. Requests are tracked per authenticated user entity reference within a fixed 1-minute window. When a user exceeds the configured limit, the server returns an HTTP 429 Too Many Requests response with a Retry-After header and a JSON error body containing a RateLimitExceeded error type.

The Developer Hub intelligent assistant applies rate limits by tier:

Expensive (default: 25 requests/min/user)
Applies to POST /v1/query (LLM inference), notebook document uploads, and Retrieval Augmented Generation (RAG) queries.
General (default: 200 requests/min/user)
Applies to all other authenticated endpoints, including conversation listing, Model Context Protocol (MCP) server management, feedback, and notebook session create, read, update, and delete operations.
Excluded (no limit)
Applies to health check endpoints (/health, /notebooks/health).

Prerequisites

  • You have deployed and configured the Developer Hub intelligent assistant instance.
  • You have administrative access to the RHDH host platform filesystem.

Procedure

  1. In your app-config.yaml file, add or modify the rateLimit block under the lightspeed section:

    lightspeed:
      rateLimit:
        expensive:
          max: 25
        general:
          max: 200
    expensive.max
    Maximum requests per minute per user for expensive endpoints such as LLM inference. Set to 0 to disable rate limiting for this tier.
    general.max

    Maximum requests per minute per user for general endpoints. Set to 0 to disable rate limiting for this tier.

    The following example shows a tightened configuration for a small deployment with limited LLM resources:

    lightspeed:
      rateLimit:
        expensive:
          max: 10
        general:
          max: 50
  2. Save the file.
  3. Restart the RHDH service to apply the updated configuration.

Verification

  1. Log in to RHDH and submit requests to the Developer Hub intelligent assistant chat interface that exceed the configured expensive limit.
  2. Confirm that the server returns an HTTP 429 Too Many Requests response with a Retry-After header indicating when to retry.

Chapter 8. Provide organization context to the Developer Hub intelligent assistant with Agent Skills

To standardize how the Developer Hub intelligent assistant answers across your teams, use Agent Skills to give it your organization’s instructions and reference material, such as development standards, troubleshooting guidance, and approved workflows.

Use Agent Skills to achieve the following goals:

Apply your development standards
Guide the Developer Hub intelligent assistant to recommend your approved libraries, patterns, and coding conventions.
Standardize troubleshooting
Give developers consistent, organization-approved answers to common errors and operational issues.
Share approved workflows
Ground responses in your operational procedures so that developers follow the same steps across teams.
Important

Developer Preview features are not supported by Red Hat in any way and are not functionally complete or production-ready. Do not use Developer Preview features for production or business-critical workloads. Developer Preview features provide early access to functionality in advance of possible inclusion in a Red Hat product offering. Customers can use these features to test functionality and provide feedback during the development process. Developer Preview features might not have any documentation, are subject to change or removal at any time, and have received limited testing. Red Hat might provide ways to submit feedback on Developer Preview features without an associated SLA.

For more information about the support scope of Red Hat Developer Preview features, see Developer Preview Support Scope.

8.1. Agent Skills for organization-specific guidance

Agent Skills enable you to provide the Developer Hub intelligent assistant with organization-specific instructions and reference material. For example, you can add skills that describe your development standards, troubleshooting procedures, or approved operational workflows.

The Developer Hub intelligent assistant loads skills from the container file system when the Lightspeed Core Service (LCORE) starts. To make skills available, you mount the skills into the lightspeed-core container and configure their location in the lightspeed-stack.yaml file.

Each skill is stored in its own directory. The directory must contain a SKILL.md file that provides the skill name, description, and instructions. A skill can also contain a references directory with supporting content. For example:

skills/
├── coding-standards/
│   ├── SKILL.md
│   └── references/
│       └── approved-libraries.md
└── openshift-troubleshooting/
    ├── SKILL.md
    └── references/
        └── common-errors.md

Skills provide instructions and reference content to the large language model (LLM). They do not modify the model.

8.2. Configure Agent Skills for Developer Hub intelligent assistant

Configure the Developer Hub intelligent assistant to load Agent Skills from the lightspeed-core container file system. Mount the skills into the container and declare their location in the lightspeed-stack.yaml file so that the skills are available when the Lightspeed Core Service (LCORE) starts.

Prerequisites

  • You have administrative access to the RHDH deployment.
  • You store your skills in a Git repository that the cluster can access, or you have the skill files available to store in a ConfigMap.

Procedure

  1. In your custom lightspeed-stack.yaml file, add the mounted directory under skills.paths:

    skills:
      paths:
        - /app-root/skills

    A path can identify either a single skill directory or a parent directory that contains multiple skill directories.

  2. Make the skills directory available to the lightspeed-core container.

    Note

    If you already mount a volume that contains your skills into the lightspeed-core container, you can skip the init-container configuration. Mount the existing volume at /app-root/skills, or update skills.paths to match its mount path.

    Otherwise, configure the volume, init container, and mount before you deploy or upgrade RHDH. This configuration ensures that the skills are present when the lightspeed-core container starts.

    1. For an Operator installation, add a strategic merge patch to the Backstage custom resource (CR):

      apiVersion: rhdh.redhat.com/v1alpha5
      kind: Backstage
      metadata:
        name: <rhdh_instance_name>
      spec:
        deployment:
          patch:
            spec:
              template:
                spec:
                  volumes:
                    - name: lightspeed-skills
                      emptyDir: {}
                  initContainers:
                    - name: fetch-lightspeed-skills
                      image: <image_with_git_client>
                      imagePullPolicy: IfNotPresent
                      command:
                        - /bin/sh
                        - -ec
                      args:
                        - |
                          git clone --depth 1 --branch <repository_ref> \
                            <skills_repository_url> /work/.source
                          test -d /work/.source/skills
                          cp -R /work/.source/skills/. /work/
                          rm -rf /work/.source
                      volumeMounts:
                        - name: lightspeed-skills
                          mountPath: /work
                  containers:
                    - name: lightspeed-core
                      volumeMounts:
                        - name: lightspeed-skills
                          mountPath: /app-root/skills
    2. For a Helm installation, add the following configuration to the values.yaml file. The chart appends these entries to the generated pod specification:

      intelligentAssistant:
        enabled: true
        core:
          extraVolumeMounts:
            - name: lightspeed-skills
              mountPath: /app-root/skills
      
      extraVolumes:
        - name: lightspeed-skills
          emptyDir: {}
      
      extraInitContainers:
        - name: fetch-lightspeed-skills
          image: <image_with_git_client>
          imagePullPolicy: IfNotPresent
          command:
            - /bin/sh
            - -ec
          args:
            - |
              git clone --depth 1 --branch <repository_ref> \
                <skills_repository_url> /work/.source
              test -d /work/.source/skills
              cp -R /work/.source/skills/. /work/
              rm -rf /work/.source
          volumeMounts:
            - name: lightspeed-skills
              mountPath: /work

      Replace the following values:

      <image_with_git_client>
      An approved image that contains Bash and Git.
      <skills_repository_url>
      The URL of the repository that contains the skills.
      <repository_ref>
      The branch or tag to retrieve.
  3. Optional: Grant a team permission to view skills information by adding the intelligent-assistant.skills permission to your RBAC policy:

    p, role:default/<team>, intelligent-assistant.skills, use, allow
    Note

    The intelligent-assistant.skills permission controls visibility only. It does not control whether LCORE loads skills or whether configured skills influence answers. Users who do not have this permission still benefit from the configured skills when they use the Developer Hub intelligent assistant. Do not use this permission as a security boundary for confidential skill names or descriptions.

  4. Apply the Backstage CR, or install or upgrade the Helm release.

    Important

    Define the configuration before the RHDH Deployment is created. When you update an existing installation, ensure that the change creates a new RHDH pod. The lightspeed-core container reads the mounted skills only during startup and does not reload changed skill files while the pod is running.

Verification

  1. Verify that the RHDH pod is ready:

    $ oc get pods -l app.kubernetes.io/instance=<rhdh_instance_name>
  2. Verify that the skill files are mounted in the lightspeed-core container:

    $ oc exec <rhdh_pod_name> -c lightspeed-core -- \
      find /app-root/skills -name SKILL.md -print
  3. Open the Developer Hub intelligent assistant and send the following query:

    List available skills.

    Verify that the LLM responds with the skills that are available in the current deployment.

  4. Ask a question that is covered by one of the configured skills, and verify that the response follows the skill instructions.

8.3. Mount Agent Skills by using a ConfigMap

If the RHDH cluster cannot access a skills repository, store each file that the skills use in a ConfigMap and mount the ConfigMap into the lightspeed-core container. Use this alternative instead of configuring an init container.

Prerequisites

  • You have administrative access to the RHDH deployment.
  • You have the skill files available on your local file system.

Procedure

  1. Create a ConfigMap that contains each skill file. Give every file a unique ConfigMap key, including files in the references directory or other supporting directories:

    $ oc create configmap <skills_config_map> \
      --from-file=coding-standards-SKILL.md=skills/coding-standards/SKILL.md \
      --from-file=coding-standards-approved-libraries.md=skills/coding-standards/references/approved-libraries.md \
      --from-file=coding-standards-example-policy.yaml=skills/coding-standards/examples/example-policy.yaml
    Note

    Specify each file individually. When you pass a directory to --from-file, oc reads only the files in the top level of that directory and does not process subdirectories recursively. As a result, files in the references directory or other subdirectories are not included.

  2. Mount the ConfigMap into the lightspeed-core container.

    1. For an Operator installation, use the same configMap volume definition in a strategic merge patch to the Backstage custom resource (CR):

      apiVersion: rhdh.redhat.com/v1alpha5
      kind: Backstage
      metadata:
        name: <rhdh_instance_name>
      spec:
        deployment:
          patch:
            spec:
              template:
                spec:
                  volumes:
                    - name: lightspeed-skills
                      configMap:
                        name: <skills_config_map>
                        items:
                          - key: coding-standards-SKILL.md
                            path: coding-standards/SKILL.md
                          - key: coding-standards-approved-libraries.md
                            path: coding-standards/references/approved-libraries.md
                          - key: coding-standards-example-policy.yaml
                            path: coding-standards/examples/example-policy.yaml
                  containers:
                    - name: lightspeed-core
                      volumeMounts:
                        - name: lightspeed-skills
                          mountPath: /app-root/skills
    2. For a Helm installation, use the following values.yaml configuration:

      intelligentAssistant:
        enabled: true
        core:
          extraVolumeMounts:
            - name: lightspeed-skills
              mountPath: /app-root/skills
      
      extraVolumes:
        - name: lightspeed-skills
          configMap:
            name: <skills_config_map>
            items:
              - key: coding-standards-SKILL.md
                path: coding-standards/SKILL.md
              - key: coding-standards-approved-libraries.md
                path: coding-standards/references/approved-libraries.md
              - key: coding-standards-example-policy.yaml
                path: coding-standards/examples/example-policy.yaml
  3. Add the mount path under skills.paths in your custom lightspeed-stack.yaml file, and apply the Backstage CR or install or upgrade the Helm release. For more information, see Configure Agent Skills.

8.4. Update Agent Skills for Developer Hub intelligent assistant

To add or update Agent Skills, update the source repository or storage volume and create a new RHDH pod. The Lightspeed Core Service (LCORE) does not detect skill changes in a running pod.

Pin the skills source to a reviewed branch, tag, or commit that follows your organization’s change-management policy. In disconnected environments, mirror the required init-container image and make the skills content available from a source that the cluster can access.

Procedure

  1. Update the skills in the source repository, ConfigMap, or storage volume.
  2. Create a new RHDH pod so that LCORE loads the updated skills. For an init-container deployment, a pod restart retrieves the configured repository reference again. For a ConfigMap deployment, update the ConfigMap before you create the new pod.

8.5. Limitations for Agent Skills

Review the following limitations before you configure Agent Skills for the Developer Hub intelligent assistant.

  • The Developer Hub intelligent assistant user interface does not provide a view that lists the available skills. You can ask the LLM to List available skills, and the LLM can respond with the skills available to the deployment.
  • Skills are loaded when the lightspeed-core container starts. Adding or modifying skill files requires a new RHDH pod.
  • Skills must be available on the local file system of the lightspeed-core container. The Developer Hub intelligent assistant does not retrieve skills directly from a remote URL.
  • Skills provide instructions and reference content only. Executable capabilities must be provided separately, for example, through Model Context Protocol (MCP) tools.
  • Skill names must be unique across all configured paths.

Chapter 9. Solve project-specific challenges with Developer Lightspeed for RHDH Notebooks

Use Developer Lightspeed for RHDH Notebooks to research, troubleshoot, and analyze projects by using a large language model (LLM) grounded in your own documentation.

Notebooks use Retrieval-Augmented Generation (RAG) so that responses are based strictly on the files you upload. Notebooks are available in overlay, docked, and fullscreen display modes. In overlay and docked modes, you can toggle the resource panel to work with notebooks on smaller screens without switching to fullscreen.

Use Notebooks to achieve the following goals:

Query your documentation
Upload project files to ask questions, summarize content, or brainstorm ideas based on those specific documents.
Troubleshoot with project-specific context
Upload project logs, architecture diagrams, or onboarding files to receive technical answers tailored to your specific environment.
Securely analyze private data
Conduct research in isolated sessions. Your uploaded data and chat history remain private and are inaccessible to other users.
Run multiple Notebooks
Uploaded documents and chat history remain available and are re-opened through the Notebook dashboard.
Verify AI responses with citations
Use the Sources chips to view the exact document excerpts used to generate an answer.
Organize research
Use metadata and tagging to categorize different research topics.

The following constraints apply during the Developer Preview:

Data boundaries
The AI can only access data within the active Notebook session.
Private access
You cannot share notebooks or documents with other team members.
Manual uploads
You must upload files directly. The tool does not support URL ingestion or web scraping.
Ephemeral defaults

Without a configured Persistent Volume (PV), all Notebook data and uploaded files are lost upon service restart.

Developer Lightspeed for RHDH Notebook fullscreen page
Important

Developer Preview features are not supported by Red Hat in any way and are not functionally complete or production-ready. Do not use Developer Preview features for production or business-critical workloads. Developer Preview features provide early access to functionality in advance of possible inclusion in a Red Hat product offering. Customers can use these features to test functionality and provide feedback during the development process. Developer Preview features might not have any documentation, are subject to change or removal at any time, and have received limited testing. Red Hat might provide ways to submit feedback on Developer Preview features without an associated SLA.

For more information about the support scope of Red Hat Developer Preview features, see Developer Preview Support Scope.

9.1. Solve project-specific challenges

Configure Red Hat Developer Hub and Red Hat Developer Lightspeed for Red Hat Developer Hub to provide users with private, document-based AI workspaces.

Prerequisites

  • A deployed instance of RHDH.
  • By using the OpenShift CLI (oc), you have access, with developer permissions, to the OpenShift Container Platform cluster aimed at containing your Developer Hub instance.
  • A Lightspeed Stack service is running and accessible to the backend.
  • A supported large language model (LLM), such as Granite 7B or higher, is available.

Procedure

  1. Enable the notebook feature and define your model by adding the following configuration to your app-config.yaml file:

    lightspeed:
      notebooks:
        enabled: true
        queryDefaults:
          model: ${NOTEBOOKS_QUERY_MODEL} # Use the exact model name
          provider_id: ${NOTEBOOKS_QUERY_PROVIDER_ID}
    Note

    If the model name is wrong, an error message is displayed in the logs and the user interface.

  2. Grant user access through role-based access control (RBAC) policies by defining permissions in your rbac-policy-csv file:

    1. Add the permission policies:

      p, role:default/<your_team_name>, intelligent-assistant.notebooks, use, allow
    2. Assign the role to specific users:

      g, user:default/<your_user_name>, role:default/<your_team_name>
  3. Apply the updated configuration and restart the service.

Verification

  1. Log in to RHDH using an account assigned to the RBAC role defined in the configuration.
  2. Confirm that the Notebooks tab is visible next to the Chat tab in the primary navigation bar.
  3. Click the Notebooks tab and ensure the My Notebooks dashboard loads without error messages.

9.2. Enable data persistence for Developer Lightspeed for RHDH Notebooks

To persist Notebook sessions, documents, and AI history across service restarts, you must configure the Notebooks storage backends to use persistent volumes.

By default, the service uses ephemeral storage in the /tmp directory, which the system clears during a pod restart.

Prerequisites

  • By using the OpenShift CLI (oc), you have access, with developer permissions, to the OpenShift Container Platform cluster aimed at containing your Developer Hub instance.
  • You have authored and provisioned a custom config map for your deployment. For more information, see link:Provision your custom Red Hat Developer Hub configuration.
  • A Persistent Volume Claim (PVC) is provisioned in your cluster and mounted to the LCORE container (for example, at /var/lib/lightspeed-data).

Procedure

  1. Update your custom config map llama-stack-configs/config.yaml file to point the kv_notebooks storage backend to your persistent mount point:

    spec:
      initContainers:
        - name: init-notebooks-dir
          # ... complete init container
      containers:
        - name: lightspeed-core
          image: quay.io/lightspeed-core/lightspeed-stack:0.5.1
          ports:
            - containerPort: 8080
          volumeMounts:
            - name: notebooks-storage
              mountPath: /var/lib/lightspeed-data
            - name: config  # ← Added all ConfigMap mounts
              mountPath: /app-root/config.yaml
              subPath: config.yaml
            - name: lightspeed-config
              mountPath: /app-root/lightspeed-stack.yaml
              subPath: lightspeed-stack.yaml
            - name: profile
              mountPath: /app-root/rhdh-profile.py
              subPath: rhdh-profile.py
          livenessProbe:  # ← Added health checks
            httpGet:
              path: /readiness
              port: 8080
          readinessProbe:
            httpGet:
              path: /readiness
              port: 8080
      volumes:  # ← Added all volume definitions
        - name: notebooks-storage
          persistentVolumeClaim:
            claimName: lightspeed-notebooks-pvc
        - name: config
          configMap:
            name: llama-stack-config
        - name: lightspeed-config
          configMap:
            name: lightspeed-core-config
        - name: profile
          configMap:
            name: rhdh-profile
  2. Update your deployment manifest to include the init container, volume mounts, and volume definitions:

    spec:
      template:
        spec:
          initContainers:
            - name: init-notebooks-storage
              image: registry.access.redhat.com/ubi9/ubi-minimal
              command: ["sh", "-c", "mkdir -p /var/lib/lightspeed-data/notebooks && chmod -R 777 /var/lib/lightspeed-data/notebooks"]
              volumeMounts:
                - name: lightspeed-notebooks
                  mountPath: /var/lib/lightspeed-data
          containers:
            - name: lightspeed-stack
              image: quay.io/lightspeed-core/lightspeed-stack:0.5.1
              ports:
                - containerPort: 8080
              volumeMounts:
                - name: lightspeed-notebooks
                  mountPath: /var/lib/lightspeed-data
                - name: config
                  mountPath: /app-root/config.yaml
                  subPath: config.yaml
              livenessProbe:
                httpGet:
                  path: /readiness
                  port: 8080
              readinessProbe:
                httpGet:
                  path: /readiness
                  port: 8080
          volumes:
            - name: lightspeed-notebooks
              persistentVolumeClaim:
                claimName: lightspeed-notebooks-pvc
            - name: config
              configMap:
                name: llama-stack-config
  3. Apply the updated configuration and restart the service.

Verification

  1. In Red Hat Developer Hub, create a Notebook and upload a test document.
  2. Send a message to the virtual assistant and verify that the response is based on the document.
  3. Restart the pod:

    $ oc delete pod <pod_name>
  4. After the pod recovers, refresh the My Notebooks dashboard.
  5. Verify that the Notebook and the uploaded file are still accessible.

Chapter 10. Get AI-assisted help for your development tasks

Use Red Hat Developer Lightspeed for Red Hat Developer Hub, a generative AI assistant in Red Hat Developer Hub (RHDH), to ask platform questions, analyze logs, generate code, and create test plans from a chat interface.

10.1. Prerequisites

  • Your platform engineer has configured the Developer Lightspeed for RHDH service in your RHDH instance.

10.2. Configure safety guards in Red Hat Developer Hub

To protect users from insecure or harmful AI model outputs, Red Hat Developer Hub (RHDH) uses Llama Guard as a default safety shield. You must configure these guards to align with your organization’s security policies.

Default safety guard configuration
The system uses Llama Guard as the default safety shield. Override these settings in the run.yaml file.
Note

The external_providers_dir parameter defaults to null and is no longer required in your configuration.

Overriding safety guards
To implement custom security layers or different safety shields, you must define a new safety provider within a custom run.yaml file.
Disabling safety guards
To run RHDH without safety guards, you must use the run-no-guard.yaml configuration file.
Important

Running without safety guards increases the risk of invalid model output. Only use this configuration in secure development environments.

Applying the no-guard configuration
To run the system without a safety guard, perform these steps:

Procedure

  1. Add the following YAML file as a config map to your namespace:

    version: 2
    image_name: redhat-ai-dev-llama-stack-no-guard
    apis:
      - agents
      - inference
      - safety
      - tool_runtime
      - vector_io
      - files
    container_image:
    external_providers_dir:
    providers:
      agents:
        - config:
            persistence:
              agent_state:
                namespace: agents
                backend: kv_default
              responses:
                table_name: responses
                backend: sql_default
          provider_id: meta-reference
          provider_type: inline::meta-reference
      inference:
        - provider_id: ${env.ENABLE_VLLM:+vllm}
          provider_type: remote::vllm
          config:
            url: ${env.VLLM_URL:=}
            api_token: ${env.VLLM_API_KEY:=}
            max_tokens: ${env.VLLM_MAX_TOKENS:=4096}
            tls_verify: ${env.VLLM_TLS_VERIFY:=true}
        - provider_id: ${env.ENABLE_OPENAI:+openai}
          provider_type: remote::openai
          config:
            api_key: ${env.OPENAI_API_KEY:=}
        - provider_id: ${env.ENABLE_VERTEX_AI:+vertexai}
          provider_type: remote::vertexai
          config:
            project: ${env.VERTEX_AI_PROJECT:=}
            location: ${env.VERTEX_AI_LOCATION:=us-central1}
        - provider_id: sentence-transformers
          provider_type: inline::sentence-transformers
          config: {}
      tool_runtime:
        - provider_id: model-context-protocol
          provider_type: remote::model-context-protocol
          config: {}
        - provider_id: rag-runtime
          provider_type: inline::rag-runtime
          config: {}
      vector_io:
        - provider_id: faiss
          provider_type: inline::faiss
          config:
            persistence:
              namespace: vector_io::faiss
              backend: faiss_kv
      files:
        - provider_id: localfs
          provider_type: inline::localfs
          config:
            storage_dir: /tmp/llama-stack-files
            metadata_store:
              table_name: files_metadata
              backend: sql_files
    storage:
      backends:
        kv_default:
          type: kv_sqlite
          db_path: /tmp/kvstore.db
        sql_default:
          type: sql_sqlite
          db_path: /tmp/sql_store.db
        sql_files:
          type: sql_sqlite
          db_path: /rag-content/vector_db/rhdh_product_docs/1.9/files_metadata.db
        faiss_kv:
          type: kv_sqlite
          db_path: /rag-content/vector_db/rhdh_product_docs/1.9/faiss_store.db
      stores:
        metadata:
          namespace: registry
          backend: faiss_kv
        inference:
          table_name: inference_store
          backend: sql_default
          max_write_queue_size: 10000
          num_writers: 4
        conversations:
          table_name: openai_conversations
          backend: sql_default
    registered_resources:
      models:
        - model_id: sentence-transformers/all-mpnet-base-v2
          metadata:
            embedding_dimension: 768
          model_type: embedding
          provider_id: sentence-transformers
          provider_model_id: /rag-content/embeddings_model
      tool_groups:
        - provider_id: rag-runtime
          toolgroup_id: builtin::rag
      vector_dbs:
        - vector_db_id: rhdh-product-docs-1_8
          embedding_model: sentence-transformers/all-mpnet-base-v2
          embedding_dimension: 768
          provider_id: faiss
    server:
      auth:
      host:
      port: 8321
      quota:
      tls_cafile:
      tls_certfile:
      tls_keyfile:
  2. Mount the config map to your Llama Stack container at /app-root/run.yaml to make sure it overrides the default image file:

    name: llama-stack
    volumeMounts:
    - mountPath: /app-root/run.yaml
      subPath: run.yaml
      name: llama-stack-config
  3. Configure the required volume:

    volumes:
    - name: llama-stack-config
      configMap:
        name: llama-stack-config

    where:

    llama-stack-config
    The config map where you added the new no-guard configuration file.
  4. Restart the deployment if it does not trigger an automatic rollout.

10.3. Best results for assistant queries

To resolve technical blockers and accelerate development tasks, you must structure your queries to give specific context to the AI assistant. Using precise prompts makes sure that Developer Lightspeed for RHDH generates relevant code snippets, architectural advice, or platform-specific instructions.

Use the following strategies to improve the accuracy of the assistant’s output during your development workflow:

Specify technologies
Instead of asking "How do I use templates?", ask "How do I create a Software Template that scaffolds a Node.js service with a CI/CD pipeline".
Give context
Include details about your environment, such as "I am deploying to OpenShift; how do I set up my catalog-info.yaml to show pod health?".
Use conversation context
Ask follow-up questions to refine an earlier answer. For example, if the assistant gives a code snippet, you can ask "Now rewrite that using TypeScript interfaces."
Validate with citations
Check the provided documentation links and citations in the response to verify that the generated advice aligns with your organization’s official standards.
Improve assistant accuracy
Rate the utility of responses by selecting the Thumbs up or Thumbs down icons. This feedback helps tune the model for your organization’s specific requirements.
Important

To keep your data secure, do not include sensitive personal information, plain text credentials, or confidential business data in your queries.

10.4. AI response monitoring and context management

Developer Lightspeed for RHDH provides features to track the AI reasoning process and keep the context of your development tasks.

Thinking cards
An expandable thinking card is displayed while the AI processes a query. A pulse animation indicates the reasoning phase. You can expand the card to view detailed reasoning or collapse it to minimize screen clutter.
Tool call transparency
An expandable card displays details for Model Context Protocol (MCP) tool calls, which you can use to monitor background processes.
Context-aware citations
Retrieval-Augmented Generation (RAG) citations appear only when the AI uses internal documentation. This makes sure that general knowledge responses remain concise.
Context preservation during model changes
When you select a different AI model, Developer Lightspeed for RHDH starts a new conversation. This keeps your earlier chats available in your history.
Structural readability
The interface formats headings and bullet points automatically to make sure responses are scannable.

10.5. Manage chats

Manage your chat history and configuration in RHDH to organize your workspace, resume earlier tasks, or find past solutions.

Chat menu options

Prerequisites

  • You have configured the Developer Lightspeed for RHDH plugin in Red Hat Developer Hub.
  • You have logged in to the portal.

Procedure

  1. Click the Open intelligent assistant floating action button at the lower right of the screen to open the chat overlay.

    Intelligent Assistant button on the Developer Hub home page
  2. Optional: Configure the interface display and server settings:

    • Click the Chatbot options icon (⋮) in the header to open the options menu.

      Chat display options
    • In the Chatbot options menu, toggle Enable pinned chats or Disable pinned chats to show or hide the pinned chats. The system enables this option by default.
    • Available only if MCP is configured: In the Chatbot options menu, click MCP settings to manage Model Context Protocol connections.
    • In the Chatbot options menu, under Display mode, select any of the following views:

      • Overlay: A floating window is displayed over the current page content.

        Lightspeed overlay display mode
      • Dock to window: A panel attaches to the right side of the screen. Activating this mode automatically closes the quick start panel if it is already open.

        Lightspeed docked mode
      • Fullscreen: A dedicated page opens for intensive chat sessions. This mode displays a revised header containing the Lightspeed logo and a horizontal tab bar, which replaces the previous main menu. Bookmark the URL in your browser to save a direct link to the chat interface.

        Lightspeed full-screen mode
  3. Start a chat or load an earlier session:

    • Enter a prompt: Type a query in the Send a message chat field and press Enter.
    • Use a sample: Click a prompt tile.
    • Change the AI model: Select a model from the model selector dropdown menu inside the prompt bar.

      Lightspeed chat AI model
    • Attach a file: Click the (+) icon on the left of the prompt bar to upload a .yaml, .json, or .txt file. Descriptive text clarifies the function of the icon.

      Attach button in the prompt bar
      1. Click the attached file name to open the Preview attachment window.
      2. View the read-only content, or click Edit to modify the file.
    • Use voice: Click the Use microphone icon. The microphone and send buttons are located on the right side of the prompt bar.
    • Control AI generation: Use the control buttons on the right side of the prompt bar. Click the Send (>) button to submit queries, or click the Stop button to halt AI generation.

      Send button in the prompt bar
      Stop button in the prompt bar during AI generation
    • Resume a chat: Select a title from the Chats list.
  4. Organize your chat history:

    • Start a new topic: Click New chat to reset the assistant’s context. When the history panel is collapsed, click the New chat icon to create a new chat.

      Lightspeed new chat
      New chat icon to start a new chat when the history panel is collapsed
    • Search history: Enter a keyword in the Search field.
    • Rename a session: Click Options next to a chat title, select Rename, enter a new name in the Rename chat? dialog, and click Rename.

      Rename chat dialog
    • Pin a chat: Click Options next to a chat title and select Pin. The chat moves to the Pinned group.
    • Sort chats: Click Sort control and choose a sorting criteria, such as Date (Newest first).

      Chat sorting options
    • Delete a chat: Click Options next to a chat title and select Delete.

      Delete chat option
    • Expand or collapse the panel: Click the Expand/collapse icon to toggle the history panel.

      Expand chat option
      Collapse chat option
      Note

      The expanded panel is resizable up to a defined maximum width.

  5. Optional: To hide the interface, if you are in the Overlay or Dock to window mode, click the Close intelligent assistant icon (X) to hide the window. If you are in Fullscreen mode, revert to the other modes and click the Close intelligent assistant icon (X). The system preserves your active query and history.
  6. Optional: In Fullscreen mode, bookmark the URL in your browser to save a direct link to the chat interface.

Verification

  1. The main window displays the active chat or selected history.
  2. The chat history list reflects renamed, pinned, or deleted entries.

10.6. Build a private knowledge base with Developer Lightspeed for RHDH Notebooks

Use Developer Lightspeed for RHDH notebooks to create isolated research environments. These workspaces allow you to analyze project data securely by using a large language model (LLM) grounded in your specific documentation.

Important

Developer Preview features are not supported by Red Hat in any way and are not functionally complete or production-ready. Do not use Developer Preview features for production or business-critical workloads. Developer Preview features provide early access to functionality in advance of possible inclusion in a Red Hat product offering. Customers can use these features to test functionality and provide feedback during the development process. Developer Preview features might not have any documentation, are subject to change or removal at any time, and have received limited testing. Red Hat might provide ways to submit feedback on Developer Preview features without an associated SLA.

For more information about the support scope of Red Hat Developer Preview features, see Developer Preview Support Scope.

10.6.1. Create isolated research workspaces

Organize your work into individual notebook sessions to keep research topics separate and private.

Notebooks are available in overlay, docked, and fullscreen display modes.

Notebook fullscreen page

Procedure

  1. In the RHDH interface, click the Open intelligent assistant floating action button (FAB) at the lower right of the screen to open the chat overlay.
  2. In your Developer Lightspeed for RHDH page, select the Notebooks tab.
  3. Click Create a new notebook to start a new workspace.

    The notebook opens with an empty resource panel. Use the display mode controls in the menu bar to switch between fullscreen, docked, and overlay modes. In overlay or docked mode, click the toggle icon to show or hide the resource panel, or click the collapse icon to minimize the current view.

    Note

    When you close a notebook, the system saves it if you edited the name or added resources. Otherwise, the system discards the unedited, empty notebook.

  4. Optional: To manage your workspaces, complete any of the following actions:

    • Rename: Click the notebook name to edit it inline. Click outside the field or press Enter to save, or press Escape to cancel. Alternatively, hover over the notebook card, click the More options icon, and select Rename.
    • Delete: Hover over the notebook card, click the More options icon, and select Delete.

      Notebook card More options menu showing Rename and Delete

Verification

  • Confirm the new notebook card appears on the My Notebooks dashboard. Each card displays the notebook name, the resource count (for example, "0 Resources", "1 Resource", or "5 Resources"), and the last updated date.

    My Notebooks dashboard showing available notebook cards

10.6.2. Provide project context to the AI

To receive answers tailored to your project, upload and manage relevant source material in your active session.

Prerequisites

  • The files that you add adhere to the following constraints:

    File limit
    You can add up to 10 files at a time. When this limit is reached, the drag-and-drop area is disabled and a tooltip indicates the limit.
    File size
    Individual files must be 25 MB or smaller.
    Notebook Capacity
    The total token count per session must not exceed 100k.
    Unsupported content
    Avoid scanned PDF images without text, audio, video, and general image files.
    Persistence requirement
    The internal SQL and KV stores must be mapped to a persistent backend to maintain the 100k token context across sessions.

Procedure

  1. Open a notebook card from the dashboard.
  2. Add resources by using one of the following methods:

    • In the sidebar, click Add (+).
    • In the main user interface, click Add a resource.
  3. In the Add resources modal, drag and drop files or click to browse. Supported formats are displayed as chips in the modal and include .txt, .log, .md, .pdf, .json, .yaml, and .yml.

    Notebook upload modal with file type chips

    If you add a file that exists in the notebook, the system highlights the duplicate and prompts you to select either Replace existing files or Ignore duplicated files.

    Notebook file exists
    Note

    A successful upload completes without a confirmation message. If an upload fails, an error notification describes the problem.

  4. Wait for the system to process and vectorize the files. This might take several seconds for larger PDFs.
  5. Optional: To manage resources in the sidebar, hover over a resource to display the More options icon. You can Rename or Delete a resource. To rename a resource, click its name to edit it inline, or select Rename from the menu.

    Resource More options menu showing Rename and Delete

Verification

  • Ensure the uploaded files appear in the Resources list in the sidebar.

10.6.3. Extract and verify document-based insights

After providing context, use the AI to perform reasoning across your files and verify the accuracy of the responses.

Lightspeed Notebook chat with sources chip

Prerequisites

  • You have uploaded documents to the active chat session to establish context.

    Note

    If a notebook contains no resources, the message bar is inactive and displays the placeholder text Ask about your resources…​. Hovering over the inactive message bar displays the tooltip Select at least one loaded resource to start chatting. You cannot enter a query until you add at least one resource.

    Procedure

    1. Enter a question in the message bar at the bottom of the screen. The message bar includes a microphone icon for voice input. The Send button appears only after you enter text. To cancel an ongoing response, click the Stop button.

      Lightspeed notebook microphone icon
    2. Analyze the response. The AI identifies relationships across all uploaded documents in the session.
    3. To verify accuracy, click the Sources chip to open a popover listing the specific documents used to generate the answer. Each source displays its filename and file type icon.

      Sources popover listing the documents used to generate the AI response
    4. Manage your workflow by using the history panel in the sidebar to expand or collapse previous interactions.

Verification

  • Confirm that the Sources popover displays the correct filenames corresponding to the AI’s response.

Chapter 11. AI model evaluation data to select the right AI model

Use the Red Hat Developer Lightspeed for Red Hat Developer Hub evaluation framework to validate the performance, accuracy, and reliability of Developer Lightspeed for RHDH.

With this automated toolset, you can measure how effectively various large language models (LLMs) answer questions based on Red Hat Developer Hub documentation.

Table 11.1. Components of the evaluation framework

ComponentDescription

Evaluation framework

Contains the core logic and scripts used to run evaluations.

Datasets

Includes the input files used to test the model.

Evaluation metrics integration

Provides scoring through various metrics, including Ragas, DeepEval, and custom metrics. Ragas is the primary metric used to validate Developer Lightspeed for RHDH performance.

11.1. Configure the evaluation environment to validate model accuracy

Set up the evaluation environment to validate the performance and accuracy of Developer Lightspeed for RHDH. Configure this evaluation to ensure the model correctly interprets documentation and provides dependable answers.

Important

Developer Preview features are not supported by Red Hat in any way and are not functionally complete or production-ready. Do not use Developer Preview features for production or business-critical workloads. Developer Preview features provide early access to functionality in advance of possible inclusion in a Red Hat product offering. Customers can use these features to test functionality and provide feedback during the development process. Developer Preview features might not have any documentation, are subject to change or removal at any time, and have received limited testing. Red Hat might provide ways to submit feedback on Developer Preview features without an associated SLA.

For more information about the support scope of Red Hat Developer Preview features, see Developer Preview Support Scope.

By performing these evaluations, you minimize the risk of the model delivering incorrect or hallucinated information to users in production.

Prerequisites

  • Install uv for Python package management (Python 3.11 or later).

Procedure

  1. Clone the evaluation repository and navigate to the directory:

    git clone https://github.com/lightspeed-core/lightspeed-evaluation
    cd lightspeed-evaluation
  2. Synchronize the environment and install dependencies:

    uv sync
  3. Configure the environment variables for the judge LLM. You can create a .env file in the root directory or export the keys directly to your terminal.

    • If you use Gemini, you must set the Gemini API key:

      export GEMINI_API_KEY="your-google-api-key"
    • If you use OpenAI, you must set the OpenAI API key:

      export OPENAI_API_KEY="your-key"
  4. Optional: If you test with a live service, set your Developer Lightspeed for RHDH service API key:

    export API_KEY="your-lightspeed-service-key"

Verification

  • Verify that the environment is synchronized and the virtual environment is active:

    uv run python --version

    The output must return Python 3.11 or later.

11.2. Prepare evaluation datasets to verify AI-generated responses

Prepare evaluation data sets to test the performance of Developer Lightspeed for RHDH. You can use pre-generated AI data sets for specific Red Hat Developer Hub releases or generate custom AI data sets from your own documentation.

Prerequisites

  • You must clone the evaluation repository to your local machine.

Procedure

  1. Download pre-generated data sets: Use this method to test the performance of specific RHDH releases. These data sets are generated using Ragas testset generation for RAG.

    1. In your terminal, navigate to the /dataset folder in the evaluation repository.
    2. Locate the .evaluation_dataset_yaml files. These files are pre-configured for the evaluation tool.
    3. To test a historical release, switch to the corresponding branch.

      For example, to access the Red Hat Developer Hub 1.8 data set, switch to the 1.8 branch.

      Important

      The main branch contains work-in-progress (WIP) data sets. Avoid using this branch for stable evaluations.

  2. Generate custom data sets: Use this method to create a new test set from your own technical documentation.

    1. Generate a diverse set of question-and-answer (Q&A) pairs by following the Ragas test data generation documentation.
    2. Ensure your Q&A pairs match the required format by reviewing the evaluation data structure configuration.

Verification

  • Verify that your custom data set matches the required schema before you start the evaluation run.

11.3. Run performance tests to ensure AI response reliability

Use the evaluation framework to run performance tests in either static mode to evaluate pre-recorded responses or dynamic mode to call a live service.

These evaluations identify performance gaps, allow you to compare different large language models (LLMs), and ensure that Developer Lightspeed for RHDH provides reliable information to users.

Prerequisites

Procedure

  1. Download the system.yaml configuration template from the repository.
  2. Configure the parameters in the system.yaml file based on your evaluation mode:

    FieldDescription

    llm

    Defines the judge LLM that scores the responses, such as gemini-2.5-pro.

    api.enabled

    Set to false for static mode to use pre-filled data. Set to true for dynamic mode to call a live service.

    api.api_base

    (Required for dynamic mode only) Provide the URL of your Developer Lightspeed for RHDH service.

    api.endpoint_type

    Specify the service configuration type: streaming or query.

  3. Execute the evaluation by using the lightspeed-eval command:

    lightspeed-eval \
      --system-config config/system.yaml \
      --eval-data config/evaluation_data.yaml \
      --output-dir ./my_evaluation_results

Verification

  • Navigate to the specified output directory and verify that the generated reports contain the model performance scores.

11.4. Analyze evaluation results to identify performance gaps

Determine the performance of Developer Lightspeed for RHDH and identify documentation areas that require model improvement by analyzing evaluation results in the repository. You can use these reports to compare performance across different large language models (LLMs) and topics.

Prerequisites

Procedure

  1. In the root of the repository, navigate to the version-specific folder within the /evaluation-result directory.
  2. Open the following files to evaluate performance:

    • Model Pass Rate: Compare the overall performance between different LLMs.
    • Topic Pass Rate: Identify performance trends and gaps within specific documentation areas.

Verification

  • Verify that the reports display data visualizations or metrics consistent with your recent evaluation run.

11.5. Evaluation metrics and historical data reference

Use the available metrics to evaluate the performance of Developer Lightspeed for RHDH at the conversation turn level.

These metrics provide a standardized way to measure the accuracy and reliability of the generated responses and the retrieved content.

MetricDescription

Faithfulness

Measures how well the answer is derived solely from the retrieved context.

Context recall

Measures whether the retrieved context contains all information required to answer the question.

Context relevance

Verifies if the retrieved documentation chunks are relevant to the user query.

Context precision without reference

Measures the ratio of useful information within the retrieved documentation chunks.

Answer correctness

Compares the generated response against the expected ground-truth response. This custom metric is implemented in the evaluation tool.

11.6. Release report and historical data

Use the latest Q&A data set and evaluation results to monitor the current performance of Developer Lightspeed for RHDH.

Access version-specific branches that contain the data sets and evaluation results required to track improvements or regressions across product releases.

Important

The main branch contains work-in-progress data for versions currently under development. For stable evaluations or historical tracking, you must switch to the branch associated with a specific release.

Release versionBranch nameData included

Latest stable

Most recent version branch

The current question and answer (Q&A) data set and evaluation results.

Historical

Previous version branches

Data sets and evaluation results for previous releases to track regressions.

Chapter 12. Appendix: LLM requirements

Review large language model (LLM) provider compatibility and system requirements for OpenAI, Red Hat OpenShift AI, vLLM, and Google Vertex AI to plan your Developer Lightspeed for RHDH infrastructure deployment.

12.1. Large language model (LLM) requirements

To plan your Developer Lightspeed for RHDH deployment, you must determine which compatible large language model (LLM) inference provider fits your infrastructure.

Developer Lightspeed for RHDH operates on a Bring Your Own Model (BYOM) architecture. Because the service does not include a native model, you must connect a compatible inference provider during installation.

The underlying LCORE service integrates with platforms that support either the OpenAI API specification or the vLLM inference engine. Developer Lightspeed for RHDH supports the following inference provider configurations:

  • OpenAI: Cloud-based inference services.
  • vLLM: Enterprise inference servers, which include models hosted on Red Hat OpenShift AI and Red Hat Enterprise Linux AI.
  • Google Vertex AI: Cloud-based inference services, which include Gemini models.

Developer Lightspeed for RHDH supports only the inference providers that the underlying LCORE service supports.

The following table lists the inference providers Developer Lightspeed for RHDH supports.

ProviderTypeSupported

OpenAI

Remote

Yes

Azure

Remote

Yes

Amazon Bedrock

Remote

Yes

Google Vertex AI

Remote

Yes

IBM watsonx

Remote

Yes

Red Hat OpenShift AI (vLLM)

Remote

Yes

Red Hat Enterprise Linux AI (RHAIIS/vLLM)

Remote

Yes

When configuring your deployment, you must account for the following provider behaviors:

  • Red Hat OpenShift AI routing: Because the configuration lacks an explicit Red Hat OpenShift AI provider option, you must route these deployments through the vLLM provider settings.
  • vLLM URL syntax: The vllm provider type communicates with endpoints that conform to the OpenAI API schema. You must manually append /v1 to the configured provider URL because the system does not add it automatically. This configuration also applies to other hosted, OpenAI-compliant inference providers.

Additional resources

12.2. OpenAI model integration for your deployment

Use OpenAI models to provide generative artificial intelligence (AI) inference services, such as GPT 5, for your Developer Lightspeed for RHDH deployment.

The system connects directly to the OpenAI API platform to route user prompts and return model insights. To configure this large language model (LLM) provider, you must have an active API key generated from your OpenAI developer account.

12.3. vLLM model integration for high-throughput inference

Use the open-source vLLM high-throughput serving framework to optimize memory utilization and manage high volumes of concurrent requests for your Developer Lightspeed for RHDH deployment.

The vLLM framework operates as an enterprise inference server that optimizes memory allocation to maximize the processing efficiency of large language models (LLMs). Integrating vLLM ensures that your environment maintains high performance and responsiveness under heavy concurrent user traffic.

Additional resources

12.4. Vertex AI integration for Gemini models

To use Gemini models with Developer Lightspeed for RHDH, you can configure Google Cloud Vertex AI to act as your managed large language model (LLM) inference provider.

The underlying LCORE service connects to Vertex AI to access hosted Gemini models. This integration provides Developer Lightspeed for RHDH with enterprise-grade language processing and chat assistance capabilities without requiring you to maintain a local inference server.

Additional resources

Chapter 13. Appendix: Manage user data security

Review data handling practices, feedback storage protocols, and model configuration architectures, such as the Bring Your Own Model approach, to evaluate and enforce information security standards for your organization.

13.1. Manage user data security

Review data routing and privacy practices to evaluate how Developer Lightspeed for RHDH handles chat messages and operational information transmitted to large language model (LLM) providers.

Developer Lightspeed for RHDH sends your chat messages directly to your configured large language model (LLM) provider. Because these messages can contain sensitive operational data regarding your cluster, users, or business environment, ensure that your provider compliance policies align with your organizational security standards.

Developer Lightspeed for RHDH has limited capabilities to filter or redact the information you submit during user interactions. To mitigate data exposure risks, do not enter proprietary or confidential information into Developer Lightspeed for RHDH. To encourage user compliance, Developer Lightspeed for RHDH displays a mandatory warning at the start of each sessions, reminding users to omit personal or sensitive details.

13.2. User feedback collection

Review how Developer Lightspeed for RHDH collects and isolates user feedback data within your cluster to manage local storage requirements and data privacy standards.

Developer Lightspeed for RHDH saves user feedback submissions, including numerical ratings and text commentary, locally within the pod filesystem. Because Red Hat does not collect, access, or transmit this data, local platform administrators must manage and monitor these storage directories.

13.3. Bring Your Own Model integration

Review Bring Your Own Model (BYOM) requirements to select and integrate an OpenAI API-compatible inference service with Lightspeed Core Service.

Developer Lightspeed for RHDH relies on a BYOM architecture that let you connect the Lightspeed Core Service (LCORE) layer to any OpenAI API-compatible inference platforms. To establish connection compatibility, your chosen inference service must satisfy the following technical criteria:

  • The service must conform to the OpenAI API specification for chat completions.
  • The host environment must match the specified infrastructure configuration and installation instructions.

Various commercial and open-source inference services support the OpenAI API specification. Because operational costs, performance metrics, and data security controls vary by provider, you must evaluate and test prospective platforms locally to select the service that best meets your organizational requirements.

Additional resources

13.4. Your compliance and data-sharing responsibility

Review compliance requirements and data-sharing responsibilities to ensure that user interactions with Developer Lightspeed for RHDH align with your organization’s data privacy policies.

All data that users submit through prompts and responses within Developer Lightspeed for RHDH is transmitted directly to your configured large language model (LLM) inference service. Platform administrators must ensure that these external data transfers comply with corporate security standards, governance frameworks, and local data protection policies.

Legal Notice

Copyright © 2026 Red Hat, Inc.
The text of and illustrations in this document are licensed by Red Hat under a Creative Commons Attribution–Share Alike 3.0 Unported license ("CC-BY-SA"). An explanation of CC-BY-SA is available at http://creativecommons.org/licenses/by-sa/3.0/. In accordance with CC-BY-SA, if you distribute this document or an adaptation of it, you must provide the URL for the original version.
Red Hat, as the licensor of this document, waives the right to enforce, and agrees not to assert, Section 4d of CC-BY-SA to the fullest extent permitted by applicable law.
Red Hat, Red Hat Enterprise Linux, the Shadowman logo, the Red Hat logo, JBoss, OpenShift, Fedora, the Infinity logo, and RHCE are trademarks of Red Hat, Inc., registered in the United States and other countries.
Linux® is the registered trademark of Linus Torvalds in the United States and other countries.
Java® is a registered trademark of Oracle and/or its affiliates.
XFS® is a trademark of Silicon Graphics International Corp. or its subsidiaries in the United States and/or other countries.
MySQL® is a registered trademark of MySQL AB in the United States, the European Union and other countries.
Node.js® is an official trademark of Joyent. Red Hat is not formally related to or endorsed by the official Joyent Node.js open source or commercial project.
The OpenStack® Word Mark and OpenStack logo are either registered trademarks/service marks or trademarks/service marks of the OpenStack Foundation, in the United States and other countries and are used with the OpenStack Foundation's permission. We are not affiliated with, endorsed or sponsored by the OpenStack Foundation, or the OpenStack community.
All other trademarks are the property of their respective owners.