diff --git a/README.md b/README.md index 2a86c347..d1693bff 100644 --- a/README.md +++ b/README.md @@ -50,41 +50,29 @@ each service component. | **Status Service** | [orpheus/status](https://github.com/fa25genai/orpheus/tree/develop/status) | | **User Interface** | [orpheus/ui](https://github.com/fa25genai/orpheus/tree/develop/ui) | - - ### API Interface Documentation -| Service | Description | OpenAPI Specification | -|----------------------------------|--------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| **Answer Generation Service** | Handles user prompts, creates lecture generation jobs, and returns a lectureId. | [Answer Generation Service](./api/answer_generation_service.yaml) | -| **Generation Status Service** | Handles the status of a lecture generation job. | [Generation Status Service](./api/generation_status_service.yaml) | -| **Content Retrieval Service** | Extracts and retrieves relevant content from instructor-provided slides and materials to support question answering. | [Content Retrieval Service](./api/content_retrieval_service.yaml) | -| **Lecture Ingestion Service** | Loads received lectures into vector database and allows deleting information related to already uploaded lecture slides. | [Lecture Ingestion Service](./api/lecture_ingestion_service.yaml) | -| **Slide Generation Service** | Generates lecture slides that conform to the layout of the respective course from a detailed lecture script. | [Slide Generation Service](./api/slide_generation_service.yaml) | -| **Slide Postprocessing Service** | Converts slides to HTML and uploads generated code to the `Generated Slide Delivery` (CDN) for distribution. | [Slide Postprocessing Service](./api/slide-postprocessing_service.yaml) | -| **Avatar Generation Service** | Produces short videos of lifelike professor avatars from a given text for the voice track with expressive narration. | [Avatar Generation Service](./api/avatar_generation_service.yaml) | -| **Video Push Service** | Uploads generated avatar videos to the `Generated Avatar Delivery` (CDN) for distribution. | TODO gather info | -| **Content Location Service** | Returns the CDN location of a slide / avatar video of a related `promptId`. | Note: not these services are not used and implemented yet, currently still relying on polling and respective status requests
[Slides Content Location Service](./api/content_location_service_slides.yaml)
[Avatar Content Location Service](./api/content_location_service_avatar.yaml) | -| **Generated Avatar Service** | Provides the generated avatar videos, retrieved by related `promptId`. | TODO gather info about CDN | -| **Generated Slide Service** | Provides the generated slides. Retrieval is done with the related `promptId`. | [Generated Slides Service](slides/delivery/README.md) | +| Service | Description | OpenAPI Specification | +|----------------------------------|--------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| **Answer Generation Service** | Handles user prompts, creates lecture generation jobs, and returns a lectureId. | [Answer Generation Service](api/answer_generation_service.yaml) | +| **Avatar Generation Service** | Produces short videos of lifelike professor avatars from a given text for the voice track with expressive narration. | Code not generated from API spec yet [Avatar Generation Service](./api/avatar_generation_service.yaml), instead the endpoint definition is found in [video.py](avatar/app/api/video.py) | +| **Content Retrieval Service** | Extracts and retrieves relevant content from instructor-provided slides and materials to support question answering. | [Content Retrieval Service](api/content_retrieval_service.yaml) | +| **Generation Status Service** | Handles the status of a lecture generation job. | [Generation Status Service](api/generation_status_service.yaml) | +| **Generated Avatar Service** | Provides the generated avatar videos, retrieved by related `promptId`. | [Generated Avatar Service](avatar/assets/README.md) | +| **Generated Slide Service** | Provides the generated slides. Retrieval is done with the related `promptId`. | [Generated Slides Service](slides/delivery/README.md) | +| **Lecture Ingestion Service** | Loads received lectures into vector database and allows deleting information related to already uploaded lecture slides. | [Lecture Ingestion Service](api/lecture_ingestion_service.yaml) | +| **Slide Generation Service** | Generates lecture slides that conform to the layout of the respective course from a detailed lecture script. | [Slide Generation Service](api/slide_generation_service.yaml) | +| **Slide Postprocessing Service** | Converts slides to HTML and uploads generated code to the `Generated Slide Delivery` (CDN) for distribution. | [Slide Postprocessing Service](api/slide_postprocessing_service.yaml) | +| **Video Push Service** | Uploads generated avatar videos to the `Generated Avatar Delivery` (CDN) for distribution. | Code not generated from API spec yet [Video Push Service](avatar/app/api/avatars.py) | ## Getting Started @@ -94,70 +82,73 @@ gather info about not yet exposed APIs (Slide Push Service, Video Push Service, The **lecturer view** can be accessed via the **Admin Button** located at the top right. -To personalize the course delivery, the lecturer is required to upload both **avatar images** and **voice samples**, followed by the relevant **course materials**. +To personalize the course delivery, the lecturer is required to upload both **avatar images** and **voice samples**, +followed by the relevant **course materials**. --- ##### 👤 Avatar Uploads +
Lecturer Avatar Upload
Three distinct avatars should be provided to represent different stages of the lecture: -- **Beginning Avatar** - - Used at the beginning of the lecture. - - Recommended: a **happy facial expression** to create a welcoming atmosphere. +- **Beginning Avatar** + - Used at the beginning of the lecture. + - Recommended: a **happy facial expression** to create a welcoming atmosphere. -- **Default/Middle Avatar** - - Used during the main lecture delivery. - - Recommended: a **neutral facial expression** to maintain focus. +- **Default/Middle Avatar** + - Used during the main lecture delivery. + - Recommended: a **neutral facial expression** to maintain focus. -- **Ending Avatar** - - Used at the end of the lecture. - - Recommended: a **happy facial expression** to close on a positive note. +- **Ending Avatar** + - Used at the end of the lecture. + - Recommended: a **happy facial expression** to close on a positive note. --- ##### 🎙️ Voice Samples +
Lecturer Audio Upload
Similarly, three voice samples should be uploaded, aligned with the same lecture stages as the avatars: -- **Beginning Voice Sample** — welcoming and engaging. -- **Default/Middle Voice Sample** — clear and neutral delivery. -- **Ending Voice Sample** — positive and encouraging tone. +- **Beginning Voice Sample** — welcoming and engaging. +- **Default/Middle Voice Sample** — clear and neutral delivery. +- **Ending Voice Sample** — positive and encouraging tone. --- ##### 📑 Course Materials +
Lecturer Material Upload
-- Upload course slides and/or pre-recorded lecture videos. -- These materials will be processed and integrated into the system by the **Document Intelligence Team**. -- For further details, refer to the [Document Intelligence README](./document-intelligence/README.md). +- Upload course slides and/or pre-recorded lecture videos. +- These materials will be processed and integrated into the system by the **Document Intelligence Team**. +- For further details, refer to the [Document Intelligence README](./document-intelligence/README.md). --- - - #### Student View -1. Choose your level of expertise by selecting a suitable character: -![alt text](StudentView_Step1.png) +1. Choose your level of expertise by selecting a suitable character: + ![alt text](StudentView_Step1.png) -2. Enter your question or choose from the predefined ones: -![alt text](StudentView_Step2.png) +2. Enter your question or choose from the predefined ones: + ![alt text](StudentView_Step2.png) -3. Wait for for the generation process. In the meantime a textual answer will be given. The lectre will start as soon as the first video is done: -![alt text](StudentView_Step3.png) +3. Wait for for the generation process. In the meantime a textual answer will be given. The lectre will start as soon as + the first video is done: + ![alt text](StudentView_Step3.png) 4. Watch the video: -![alt text](StudentView_Step4.png) + ![alt text](StudentView_Step4.png) ### Development Setup @@ -167,12 +158,23 @@ Similarly, three voice samples should be uploaded, aligned with the same lecture ```bash cp exampleEnv .env ``` -2. Make sure to supply values for at least one AI-Model (e.g. AWS) and define the respective model names (`MODEL_NAME`, `SPLITTING_MODEL`, `SLIDESGEN_MODEL`) - 1. How to get AWS Keys? - 1. You need an AWS Hackathon Account - 2. Go to https://slalom-hackathon.awsapps.com/start/#/?tab=accounts - 3. Go to slalom_IsbUsersPS - 4. ... +2. Make sure to supply values for at least one AI-Model (e.g. AWS) and define the respective model names (`MODEL_NAME`, + `SPLITTING_MODEL`, `SLIDESGEN_MODEL`) + 1. How to get AWS Keys? (with Hackathon license) + 1. You need an AWS Hackathon Account + 2. Login at https://slalom-hackathon.awsapps.com/start/#/?tab=accounts + 3. Go to [slalom_IsbUsersPS](https://eu-central-1.console.aws.amazon.com/console/home?region=eu-central-1#) + 4. Go to [Amazon Bedrock](https://eu-central-1.console.aws.amazon.com/bedrock/home?region=eu-central-1#) + 5. [View API keys](https://eu-central-1.console.aws.amazon.com/bedrock/home?region=eu-central-1#/api-keys?tab=short-term) + 6. Click `Generate short-term API keys` (`Generate long-term API keys` did not work with our license) + 7. Copy the upper `API Key` that starts with `bedrock-api-key...` + 2. Update the model (AWS includes some user specific tag there, so API keys of other people won't work) + 1. You need an AWS Hackathon Account + 2. Login at https://slalom-hackathon.awsapps.com/start/#/?tab=accounts + 3. Go to [slalom_IsbUsersPS](https://eu-central-1.console.aws.amazon.com/console/home?region=eu-central-1#) + 4. Go to [Amazon Bedrock](https://eu-central-1.console.aws.amazon.com/bedrock/home?region=eu-central-1#) + 5. [Infer -> Cross-region inference](https://eu-central-1.console.aws.amazon.com/bedrock/home?region=eu-central-1#/inference-profiles) + 6. Choose the model that you want to use, copy the `Inference profile ARN` that starts with `arn:aws:bedrock:...` 3. You can overwrite the global `.env` file values with service specific `.env` files #### 1. Install Python 3.13.7 using pyenv @@ -263,37 +265,51 @@ Expected output: Python 3.13.7 ```powershell Invoke-WebRequest -UseBasicParsing -Uri "https://raw.githubusercontent.com/pyenv-win/pyenv-win/master/pyenv-win/install-pyenv-win.ps1" -OutFile "./install-pyenv-win.ps1"; &"./install-pyenv-win.ps1"; Remove-Item "./install-pyenv-win.ps1" ``` +
Troubleshooting Common Installation Issues ### 1. Script Execution is Disabled -- **Issue:** You receive an error in PowerShell stating `...cannot be loaded because running scripts is disabled on this system.` -- **What to do:** This is due to PowerShell's Execution Policy. Run PowerShell as **Administrator** and execute the following command to allow the script to run for the current session, then try the installation command again. - ```powershell - Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope Process - ``` + +- **Issue:** You receive an error in PowerShell stating + `...cannot be loaded because running scripts is disabled on this system.` +- **What to do:** This is due to PowerShell's Execution Policy. Run PowerShell as **Administrator** and execute the + following command to allow the script to run for the current session, then try the installation command again. + ```powershell + Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope Process + ``` ### 2. `pyenv` Command Not Found After Installation -- **Issue:** After the installer finishes, opening a new terminal and typing `pyenv` results in a `command not found` error. -- **What to do:** The installer couldn't modify your User `PATH` environment variable correctly, or your terminal session needs to be refreshed. - 1. **Restart your terminal:** Close and reopen PowerShell/CMD completely. - 2. **Restart your computer:** A full restart will ensure environment variables are reloaded. - 3. **Manually add to PATH:** If it still fails, you must add the following two paths to your User `PATH` environment variables. + +- **Issue:** After the installer finishes, opening a new terminal and typing `pyenv` results in a `command not found` + error. +- **What to do:** The installer couldn't modify your User `PATH` environment variable correctly, or your terminal + session needs to be refreshed. + 1. **Restart your terminal:** Close and reopen PowerShell/CMD completely. + 2. **Restart your computer:** A full restart will ensure environment variables are reloaded. + 3. **Manually add to PATH:** If it still fails, you must add the following two paths to your User `PATH` environment + variables. - `%USERPROFILE%\.pyenv\pyenv-win\bin` - `%USERPROFILE%\.pyenv\pyenv-win\shims` ### 3. System Python Overrides `pyenv` Version -- **Issue:** You've set a Python version with `pyenv global` or `pyenv local`, but running `python --version` shows your old system version (or opens the Microsoft Store). -- **What to do:** This is a `PATH` priority issue. - 1. **Disable Windows App Execution Aliases:** Go to `Start > Manage App Execution Aliases` and turn **off** the aliases for `python.exe` and `python3.exe`. This is the most common cause. - 2. **Check your `PATH` order:** Ensure the `pyenv` `shims` and `bin` paths appear *before* any other Python installation paths in your environment variables. + +- **Issue:** You've set a Python version with `pyenv global` or `pyenv local`, but running `python --version` shows your + old system version (or opens the Microsoft Store). +- **What to do:** This is a `PATH` priority issue. + 1. **Disable Windows App Execution Aliases:** Go to `Start > Manage App Execution Aliases` and turn **off** the + aliases for `python.exe` and `python3.exe`. This is the most common cause. + 2. **Check your `PATH` order:** Ensure the `pyenv` `shims` and `bin` paths appear *before* any other Python + installation paths in your environment variables. ### 4. Shims Are Not Working for New Packages -- **Issue:** You install a package with a command-line tool (like `pipx` or `poetry`) using `pip`, but the command isn't available in your terminal. -- **What to do:** You need to rebuild the `pyenv` shims so it's aware of the new executable. - ```powershell - pyenv rehash - ``` + +- **Issue:** You install a package with a command-line tool (like `pipx` or `poetry`) using `pip`, but the command isn't + available in your terminal. +- **What to do:** You need to rebuild the `pyenv` shims so it's aware of the new executable. + ```powershell + pyenv rehash + ```
2. Add pyenv to your PowerShell session @@ -321,11 +337,13 @@ Expected output: Python 3.13.7 ``` Expected output: Python 3.13.7 - (Optional) Check which version pyenv is managing - ```powershell - pyenv version - ``` - Expected output: 3.13.7 (set by C:\Users\YourUser\.pyenv\pyenv-win\version) +(Optional) Check which version pyenv is managing + +```powershell +pyenv version +``` + +Expected output: 3.13.7 (set by C:\Users\YourUser\.pyenv\pyenv-win\version) #### 2. Install Poetry 2.2.1 diff --git a/api/content_location_service_avatar.yaml b/api/content_location_service_avatar.yaml deleted file mode 100644 index 80887c8f..00000000 --- a/api/content_location_service_avatar.yaml +++ /dev/null @@ -1,119 +0,0 @@ -openapi: 3.1.0 -info: - title: Avatar Content Location Service API - version: "0.1" - description: | - API for the Orpheus video generation. - From the repository: "The Orpheus System transforms static slides into interactive lecture videos with lifelike professor avatars, combining expressive narration, visual presence, and dynamic content to create engaging, personalized learning experiences." - license: - name: MIT - url: "https://opensource.org/licenses/MIT" - termsOfService: "https://github.com/fa25genai/orpheus" - contact: - name: "Orpheus Project" - url: "https://github.com/fa25genai/orpheus/issues/new" - -servers: - - url: "http://localhost:8080" - description: "Local development" - -tags: - - name: video - description: "Endpoints for video generation, retrieval and status" - -paths: - - /v1/video/{promptId}/getLocation: - get: - tags: - - slides - summary: "Get the URL for the avatar video that has been generated (or is currently generated)." - operationId: getGenerationStatus - parameters: - - name: promptId - in: path - required: true - schema: - type: string - format: uuid - description: "The promptId returned by /v1/slides/generate" - responses: - "200": - description: "URL to the CDN where the slides are stored (response might take a while if the slides are still generated)" - content: - application/json: - schema: - $ref: "#/components/schemas/AvatarContentLocationResponse" - examples: - done: - value: - promptId: "01997fe4-864e-7da0-b959-b44e925a3c6e" - slidesUrl: "https://cdn.example.com/videos/0f2f9d77.mp4" - "404": - description: "Slides were not found (wrong promptId?)" - content: - application/json: - schema: - $ref: "#/components/schemas/Error" - -components: - schemas: - UserProfile: - type: object - description: "Information about the user and their preferences." - required: - - id - - role - - language - properties: - id: - type: string - description: "Unique identifier of the user." - role: - type: string - enum: [ student, instructor ] - description: "User's role." - language: - type: string - enum: [ german, english ] - description: "Primary language." - preferences: - type: object - description: "Generation/presentation preferences." - properties: - answerLength: - type: string - enum: [ short, medium, long ] - languageLevel: - type: string - enum: [ basic, intermediate, advanced ] - expertiseLevel: - type: string - enum: [ beginner, intermediate, advanced, expert ] - includePictures: - type: string - enum: [ none, few, many ] - enrolledCourses: - type: array - description: "Course IDs the user is enrolled in." - items: - type: string - - AvatarContentLocationResponse: - type: object - properties: - lectureId: { type: string, format: uuid } - resultUrl: { type: string, format: uri, description: "URL as soon as it is ready" } - error: { $ref: "#/components/schemas/Error" } - Error: - type: object - properties: - code: - type: string - message: - type: string - -externalDocs: - description: "Project repository and README" - url: "https://github.com/fa25genai/orpheus" - diff --git a/api/content_location_service_slides.yaml b/api/content_location_service_slides.yaml deleted file mode 100644 index 2594388a..00000000 --- a/api/content_location_service_slides.yaml +++ /dev/null @@ -1,130 +0,0 @@ -openapi: 3.1.0 -info: - title: Slides Content Location Service API - version: "0.1.0" - description: | - API for the Orpheus slide generation. - From the repository: "The Orpheus System transforms static slides into interactive lecture videos with lifelike professor avatars, combining expressive narration, visual presence, and dynamic content to create engaging, personalized learning experiences." - License: MIT (see repository). - termsOfService: "https://github.com/fa25genai/orpheus" - contact: - name: "Orpheus Project" - url: "https://github.com/fa25genai/orpheus/issues/new" - license: - name: MIT - url: "https://opensource.org/licenses/MIT" - -servers: - - url: "http://localhost:30606" - description: "Local development (random port chosen: 30606)" - - url: "http://orpheus-service-slidegen:8080" - description: "DNS service name for in-cluster service" - -tags: - - name: slides - description: "Endpoints for slide generation, retrieval and status" - -paths: - - /v1/slides/{promptId}/getLocation: - get: - tags: - - slides - summary: "Get the URL for slides that have been generated (or are currently generated)." - operationId: getGenerationStatus - parameters: - - name: promptId - in: path - required: true - schema: - type: string - format: uuid - description: "The promptId returned by /v1/slides/generate" - responses: - "200": - description: "URL to the CDN where the slides are stored (response might take a while if the slides are still generated)" - content: - application/json: - schema: - $ref: "#/components/schemas/GenerationStatusResponse" - examples: - done: - value: - promptId: "01997fe4-864e-7da0-b959-b44e925a3c6e" - slidesUrl: "https://someCDNUrl" - "404": - description: "Slides were not found (wrong promptId?)" - content: - application/json: - schema: - $ref: "#/components/schemas/Error" - -components: - schemas: - UserProfile: - type: object - description: "Schemaless additional information about the user (e.g. preferences regarding slide style)." - - SlideItem: - type: object - description: "One slide entry in the structure summary" - properties: - content: - description: Human readable text, describing which lecture contents will be contained in the slide - type: string - - SlideStructure: - type: object - description: "High-level structure of the slide deck returned early for UI and navigation" - properties: - pages: - type: array - items: - $ref: "#/components/schemas/SlideItem" - example: - pages: - - content: "Title slide introducing the topic \"for loops\"" - - content: "Loops are a programming structure which allows to repeatedly execute the same code." - - content: "Simple example of a for loop with the text \"...\"" - - GenerationAcceptedResponse: - type: object - description: "Returned immediately after generation request accepted" - properties: - promptId: - type: string - format: uuid - status: - type: string - enum: [ IN_PROGRESS, FAILED, DONE ] - createdAt: - type: string - format: date-time - structure: - $ref: "#/components/schemas/SlideStructure" - description: "Structure preview for progressing other generation steps" - - GenerationStatusResponse: - type: object - properties: - promptId: - type: string - format: uuid - slidesUrl: - type: string - format: url - error: - $ref: "#/components/schemas/Error" - - Error: - type: object - properties: - code: - type: string - message: - type: string - -externalDocs: - description: "Project repository and README" - url: "https://github.com/fa25genai/orpheus" - diff --git a/avatar/assets/README.md b/avatar/assets/README.md index 29474c00..d66220ed 100644 --- a/avatar/assets/README.md +++ b/avatar/assets/README.md @@ -1,8 +1,10 @@ -# Video Assets (NGINX) +# **Generated Avatar Service** This service is a lightweight NGINX container that serves MP4 files produced by the Python service. Files are written to a shared Docker volume and exposed over HTTP at `/videos/.mp4`. +If the service is started via the [top level docker compose](../../docker-compose.yaml) you can access the videos via `http://localhost:3000/videos/jobs//.mp4` + ## Structure ``` assets/ diff --git a/slides/delivery/README.md b/slides/delivery/README.md index 1c245bed..15669fb4 100644 --- a/slides/delivery/README.md +++ b/slides/delivery/README.md @@ -1,4 +1,4 @@ -# Orpheus **Generated Slides Service** +# **Generated Slides Service** This component serves previously generated slidesets to use in the frontend.