Content Indexing (Open Beta)

Content Indexing (Open Beta)

The more content Learn Amp can read, the better search works — for everyone. This page explains what gets indexed, how different formats are processed, and what it means for learners trying to find the right content at the right moment.


Overview

Learn Amp uses two complementary search systems, and both benefit from the same indexed content:

  • Standard in-app search — available to all users across all roles

  • Volt AI search (powered by OpenSearch with vector embeddings) — available to Owners and Admins via Volt

When content is indexed, it becomes searchable across both systems.


What Gets Indexed

Content Types and Fields

Content Type

Fields Indexed

Notes

Content Type

Fields Indexed

Notes

Items

Title, description, tags, skills, body content

Core content fields

Items (audio/video)

AI-generated transcript

Requires AI Media Transcriptions feature to be enabled

Items (PDF)

Extracted text content

Requires PDF Text Extraction feature to be enabled

Items (SCORM/eLearning)

Extracted text content + embedded audio/video transcripts

Requires Volt AI Content Indexing to be enabled

Learnlists

Title, description, tags, skills

Carousels

Title, description

Quizzes

Title, description, content

Volt AI search only

💡 Transcript, PDF, and SCORM text extraction means users can find content based on what's said or written inside a file — not just its title or description.

Audio and Video Transcriptions

When AI Media Transcriptions are enabled for your organisation, Learn Amp automatically submits audio and video items to a transcription service. The resulting transcript is stored against the item and included in both the standard search index and the Volt AI search index.

This means a learner searching for a phrase spoken in a video can find that item through search — even if the title or description doesn't mention those words.

Items that had captions added prior to this feature being enabled have also been backfilled, so existing transcribed content is already searchable.

PDF Text Extraction

When PDF Text Extraction is enabled for your organisation, Learn Amp extracts the text content from PDF items and includes it in the search index. This is standard text extraction — it only works with digitally created PDFs that contain selectable text.

⚠️ PDFs that consist entirely of scanned images (without an underlying text layer) cannot be indexed, as this process does not use optical character recognition (OCR).

SCORM Content Extraction

When Volt AI Content Indexing is enabled for your organisation, Learn Amp automatically extracts educational text content from SCORM packages and includes it in both the standard search index and the Volt AI search index.

Learn Amp identifies the authoring tool behind each package and picks the most effective extraction strategy for it — so you get the best possible results without any manual configuration. The text is then cleaned and normalised to strip out navigation labels, boilerplate, and formatting noise — leaving only the content that's actually worth searching.

💡 When you first enable Volt AI Content Indexing, Learn Amp automatically processes all your existing SCORM packages — so your full library becomes searchable straight away.

What gets extracted:

  • Slide and page text content

  • Quiz questions and answers

  • Definitions and glossary terms

  • Case study content

  • Transcripts of embedded audio and video files (transcribed automatically as part of the extraction process — no separate AI Media Transcriptions enablement required)

Supported authoring tools:

  • Easygenerator

  • Articulate Rise

  • Articulate Storyline

  • Adobe Captivate

  • iSpring

  • Evolve/Adapt

  • Generic packages (fallback extraction is applied for unrecognised tools)

Limitations to be aware of:

  • Only SCORM packages are supported — xAPI and AICC packages are not currently supported for content extraction

  • Dispatch and wrapper packages (such as Go1 Dispatch or Rustici Cross-Domain) redirect to external content, so only the item title is indexed for those

  • A maximum of 10 embedded audio or video files per package are transcribed

  • Embedded media files must be between 1 MB and 500 MB to be eligible for transcription

  • The maximum supported SCORM package (ZIP) size is 2 GB

⚠️ Scanned images embedded within SCORM packages cannot be indexed — the extraction process does not use OCR.


How Volt AI Search Uses This Content

Volt's AI search uses vector embeddings to find semantically relevant content — meaning it can surface items that are conceptually related to a query, even when the wording doesn't match exactly.

To understand what each item is about, Volt looks at the following fields:

  • Title and name

  • Description and short description

  • Extracted content text (from PDFs and SCORM packages)

  • AI-generated transcript

  • Tags and skills

  • Body content

The richer the content, the more accurately Volt can match it to learner queries. Adding transcriptions, PDF text, and SCORM content significantly expands Volt's ability to surface relevant material.


How AI Is Used in Content Indexing

Content indexing uses AI at two stages: extracting and enriching content, and making it searchable. Here's what's involved:

Stage

AI Service

What It Does

Stage

AI Service

What It Does

Audio and video transcription

AssemblyAI (Universal-2 model)

Converts spoken audio and video content into text — used for both standalone media items and audio/video embedded within SCORM packages

SCORM text cleaning

OpenAI (GPT-4o-mini)

Cleans raw text extracted from SCORM packages — removes navigation labels, UI text, and formatting noise while preserving educational content

Volt AI search embeddings

OpenAI (text-embedding-3-small)

Converts indexed text into vector embeddings, enabling Volt to find semantically relevant content

All AI services are accessed via EU-hosted endpoints. None of these services use your content to train their models — see the data privacy FAQ below and the Security, Privacy and Transparency article for full details.

💡 PDF text extraction does not use AI — it extracts selectable text directly from PDF files using standard text processing.


Feature Enablement

Your content enrichment features are controlled at company level — here's what's available and how to switch each one on:

Feature

How It's Enabled

Who Benefits

Feature

How It's Enabled

Who Benefits

AI Media Transcriptions

Settings > AI Features (toggled by Owner/Admin)

All users (standard search) + Volt users

PDF Text Extraction

Enabled by Learn Amp support team

All users (standard search) + Volt users

SCORM Content Indexing

Volt AI Content Indexing consent on Settings > AI Features

All users (standard search) + Volt users

To request PDF Text Extraction be enabled for your organisation, contact your Customer Success Manager or reach out to Learn Amp support.


Pre-requisites

Role Requirements

Action

Required Role

Action

Required Role

Enable AI Media Transcriptions

Owner or Admin

Enable SCORM Content Indexing

Owner or Admin

Access Volt AI search

Owner or Admin

Benefit from indexed content in standard search

All roles


FAQs

Q: Are older items with transcriptions also searchable?
Yes. Items that already had captions or transcriptions before this feature was fully rolled out have been backfilled, so they are included in the search index.

Q: Can Volt search inside the content of a PDF?
Yes, provided PDF Text Extraction is enabled for your organisation and the PDF contains selectable text. Scanned image-only PDFs are not currently supported.

Q: Does this affect eLearning (SCORM) content?
Yes — when Volt AI Content Indexing is enabled, Learn Amp extracts educational text from SCORM packages (including quiz questions, slide content, and definitions) and transcribes any embedded audio or video. This content is then indexed for both standard search and Volt AI search.

Q: Which SCORM authoring tools are supported?
Learn Amp currently supports content extraction from Easygenerator, Articulate Rise, Articulate Storyline, Adobe Captivate, iSpring, and the Evolve/Adapt framework. For packages built with other tools, a generic fallback extraction method is applied, which captures less structured content than tool-specific strategies — but it will still index plain text throughout the package, so it's worth enabling.

Q: What about xAPI or AICC packages?
Only SCORM packages are currently supported for content extraction. xAPI and AICC packages are not currently supported. Standard item fields (title, description, tags) are still indexed for all item types regardless of format.

Q: Is my content used to train AI models?
No. Content sent to AI services during indexing is used solely to process your data and return results — it is never used to train or improve any AI models. Learn Amp has zero data retention and no-training agreements in place with all AI providers used in content indexing (OpenAI and AssemblyAI), meaning your content is processed and immediately discarded by those services. For full details on how Volt handles data, see Security, Privacy and Transparency.

Q: What happens if I turn off Volt AI Content Indexing or archive content?
When you archive or delete an item, its indexed data (including any vector embeddings used by Volt AI search) is removed from the search index straight away. If you turn off Volt AI Content Indexing at company level, no new content will be indexed and Volt AI search will stop returning results — your existing indexed data remains stored but is no longer queryable. Turning indexing back on will resume normal indexing without needing to reprocess your full library.

Q: Why can't I find content I know exists?
A few things to check: confirm the relevant feature is enabled for your organisation; check that the item has been published; and verify the PDF contains selectable text (rather than being a scanned image). For SCORM content, check that the package is not a dispatch or wrapper package — these only index the title. If the item was uploaded very recently, wait a few minutes for indexing to complete.


Troubleshooting

Symptom

Possible Cause

What to Do

Symptom

Possible Cause

What to Do

Video or audio item doesn't appear in search results for spoken content

AI Media Transcriptions not enabled, or transcript still processing

Check Settings > AI Features; allow time for transcription to complete on recently uploaded items

PDF content not appearing in search

PDF Text Extraction not enabled, or PDF is image-only

Contact your CSM to confirm the feature is active; verify the PDF has a text layer

SCORM content not appearing in search results

SCORM Content Indexing not enabled, or package is still being processed

Enable Volt AI Content Indexing on Settings > AI Features; allow time for indexing to complete on recently uploaded packages

SCORM dispatch or wrapper package returns no content in search

Dispatch packages redirect to external content — only the title is indexed

This is expected behaviour for dispatch packages; content lives externally and cannot be extracted

SCORM embedded media not being transcribed

More than 10 media files in the package, or file size outside supported range

Only the first 10 embedded media files per package are transcribed; files must be between 1 MB and 500 MB

Volt returns no results for a topic you know is covered

Content may not yet be indexed, or embeddings are still being generated

Allow a few minutes after upload; confirm the item is published

Standard search results differ from Volt results

The two systems use different matching logic

This is expected — standard search uses keyword matching, Volt uses semantic vector search