An interactive AI art installation that transforms human emotion into poetry, narration, and generative visuals.
Magic Mirror was created for UTEC's Art and Technology course and exhibited at the Museum of Contemporary Art of Lima (MAC Lima). The project explores a simple question: what happens when a machine interprets an emotion and reflects it back as art?
I designed and built the software integration behind the installation, connecting real-time computer vision, an emotion-recognition model, generative AI, text-to-speech, and a projected Pygame experience.
- A visitor raises both open palms to begin the interaction.
- The camera captures facial input and the emotion model produces a probabilistic classification.
- The detected emotion becomes context for a poem generated through OpenAI.
- Google Text-to-Speech narrates the poem while a matching Deforum/Stable Diffusion visual is presented.
- The resulting media becomes part of the installation's evolving digital mural.
flowchart LR
A[Camera input] --> B[Hand-gesture trigger]
B --> C[Emotion recognition]
C --> D[Poetry generation]
C --> E[Visual selection]
D --> F[Text-to-speech]
E --> G[Pygame installation]
F --> G
G --> H[Digital mural]
- Coordinated camera input, gesture detection, emotion inference, media playback, and interface state in a single interactive flow.
- Integrated an OpenAI language model to turn model output into short-form poetry.
- Added text-to-speech narration and synchronized the generated poem with artistic video.
- Connected pre-generated Deforum/Stable Diffusion media with the runtime experience.
- Designed the interaction for a public exhibition rather than a conventional desktop application.
| Area | Tools |
|---|---|
| Interaction | Python, Pygame, camera input |
| Computer vision | MediaPipe, hand detection, _mini_XCEPTION emotion model |
| Generative AI | OpenAI API, Deforum Stable Diffusion |
| Audio | Google Text-to-Speech |
| Presentation | Jinja2, GitHub Pages, Firebase-hosted media |
Magic Mirror was built as a site-specific installation with a known camera, display, operating environment, and curated media library. That shaped several decisions:
- Latency and continuity mattered more than supporting arbitrary hardware.
- The visual library was generated before the exhibition so the live experience remained responsive.
- Gesture activation provided a physical, discoverable way to start the installation.
- The emotion classifier was used as an artistic input, not as a clinical or psychological assessment.
This repository preserves the original 2023 exhibition software and media. It is an installation archive, not a maintained production service or a one-command deployment.
The original environment used Python 3.8, camera hardware, a local emotion-model asset, generated media, and third-party API configuration. Modern package versions or different hardware may require adaptation. No current reliability, model-accuracy, or cross-platform guarantees are implied.
The completed installation was presented to the public at MAC Lima. The project is meaningful to me because it combined applied AI with physical interaction and made a technical pipeline understandable through an immediate human experience.
For professional context or questions about the engineering decisions, connect with me on LinkedIn.