---
title: Local Real-Time Subtitles on Ubuntu with Moonshine and Pympress
tags:  #ubuntu  
author: [Mitsuya Shibata](https://image.docswell.com/user/mtyshibata)
site: [Docswell](https://www.docswell.com/)
thumbnail: https://bcdn.docswell.com/page/37K94WK97D.jpg?width=480
description: Language barriers remain a significant challenge in international conferences. Not all speakers are fluent in English, and even when they are, their speech may not always be easy for diverse audiences to understand. While some events provide professional captioning or simultaneous interpretation, such services are not always available for every session.  In this talk, I will present a self-contained approach that enables speakers to generate real-time subtitles directly from their own laptop. By combining Moonshine Voice, a lightweight and efficient speech-to-text (STT) system that can run even on devices like Raspberry Pi, with Pympress, a PDF-based presentation tool, we can build a fully local, automatic subtitle system on Ubuntu.  This approach allows speakers to make their talks more accessible without relying on external infrastructure or services, making it easier to participate in international events regardless of location or available resources.  https://events.canonical.com/event/146/contributions/935/
published: August 09, 26
canonical: https://image.docswell.com/s/mtyshibata/K271DG-UbuConAsia2026
---
# Page. 1

![Page Image](https://bcdn.docswell.com/page/37K94WK97D.jpg)

Local Real-Time Subtitles on Ubuntu
with Moonshine and Pympress
Mitsuya Shibata
Ubuntu Japanese Team / SOUM Corporation
2026-08-09


# Page. 2

![Page Image](https://bcdn.docswell.com/page/LJ3WL1Z9J5.jpg)

Introduction


# Page. 3

![Page Image](https://bcdn.docswell.com/page/8JDKNXR8EG.jpg)

Two Questions
• Can you pronounce English well?
• Can you easily understand spoken English?
First, I’d like to ask you two questions.


# Page. 4

![Page Image](https://bcdn.docswell.com/page/VEPK5PWW78.jpg)

Two Questions
• Can you pronounce English well?
• Can you easily understand spoken English?
If you’re attending UbuCon Asia and COSCUP, you can probably communicate in English.


# Page. 5

![Page Image](https://bcdn.docswell.com/page/27VVP2897Q.jpg)

Two Questions
• Can you pronounce English well?
• Can you easily understand spoken English?
But I can’t!
If you’re attending UbuCon Asia and COSCUP, you can probably communicate in English.


# Page. 6

![Page Image](https://bcdn.docswell.com/page/5JGLPR547L.jpg)

Today’s topics
That’s all.


# Page. 7

![Page Image](https://bcdn.docswell.com/page/47QY2VZREP.jpg)

Today’s topics
That’s all.
Thank you very much!


# Page. 8

![Page Image](https://bcdn.docswell.com/page/KE4WLM3QJ1.jpg)

The English Cliff
International Tech Community Summit
(UbuCon Asia, COSCUP)
(・∀・) 人 (≧▽≦)
------------------------+
|
|
|
The English Cliff |
|
|
| (´・ω ・`) &lt; &quot;I can&#039;t speak English...&quot;
+------------------------------------------At international technical communities like UbuCon Asia, English is the common language. There
is a language barrier.


# Page. 9

![Page Image](https://bcdn.docswell.com/page/L71YL81LJG.jpg)

The English Cliff
International Tech Community Summit
(UbuCon Asia, COSCUP)
(・∀・) 人 (≧▽≦)
------------------------+
|
|
|
The English Cliff |
|
|
| (´・ω ・`) &lt; &quot;I can&#039;t speak English...&quot;
+------------------------------------------If you can’t speak or understand English, it’s hard to communicate in person.


# Page. 10

![Page Image](https://bcdn.docswell.com/page/G7WGVZ8QE2.jpg)

The English Cliff
International Tech Community Summit
(UbuCon Asia, COSCUP)
(・∀・) 人 (≧▽≦)
------------------------+
|
|
|
The English Cliff |
|
|
| (´・ω ・`) &lt; &quot;I can&#039;t speak English...&quot;
+------------------------------------------So I started wondering: Could open source help us overcome this barrier?


# Page. 11

![Page Image](https://bcdn.docswell.com/page/4JZL4183E3.jpg)

Actually...
It’s not just English.
To be honest, it isn’t only about English.


# Page. 12

![Page Image](https://bcdn.docswell.com/page/YE6W3LPZEV.jpg)

Actually...
It’s not just English.
I’m not even very good at communicating in my native language.
I’m not very good at communicating even in my native language.


# Page. 13

![Page Image](https://bcdn.docswell.com/page/GE5ML1K9E4.jpg)

Actually...
It’s not just English.
I’m not even very good at communicating in my native language.
I’m not even good at saying “hello”.


# Page. 14

![Page Image](https://bcdn.docswell.com/page/9729L1W5JR.jpg)

The Real Problem
International Tech Community Summit
(UbuCon Asia, COSCUP)
(・∀・) 人 (≧▽≦)
------------------------+
|
The English Cliff |
| (´・ω ・`) &lt; &quot;I can&#039;t speak English...&quot;
+---------------------+
|
The Confidence Cliff |
| Me &lt; &quot;H-h-hello...&quot;
+--------------------So this is the real problem.


# Page. 15

![Page Image](https://bcdn.docswell.com/page/DJY4QZL67M.jpg)

The Real Problem
International Tech Community Summit
(UbuCon Asia, COSCUP)
(・∀・) 人 (≧▽≦)
------------------------+
|
The English Cliff |
| (´・ω ・`) &lt; &quot;I can&#039;t speak English...&quot;
+---------------------+
|
The Confidence Cliff |
| Me &lt; &quot;H-h-hello...&quot;
+--------------------I have to overcome not only the English Cliff, but also the Confidence Cliff.


# Page. 16

![Page Image](https://bcdn.docswell.com/page/V7NYK3L1E8.jpg)

But...


# Page. 17

![Page Image](https://bcdn.docswell.com/page/YJ9PV94Y73.jpg)

But...
I still wanted to speak at UbuCon Asia.


# Page. 18

![Page Image](https://bcdn.docswell.com/page/GJ8DP9QKJD.jpg)

But...
I still wanted to speak at UbuCon Asia.
That’s how this project began. I wanted to find a way to give a talk in English.


# Page. 19

![Page Image](https://bcdn.docswell.com/page/LJLM1WXPER.jpg)

Commercial Break
• Ubuntu Japanese LoCo Team member
• Embedded software engineer at SOUM
Corporation1
• Author of Ubuntu-related articles in Jpanese
• 「はじめての Ubuntu」2
• Ubuntu Weekly Recipe3
• Ubuntu 日和4
• Software Design5
• 日経 Linux6
1
https://www.soum.co.jp/
https://gihyo.jp/book/2025/978-4-297-15277-2
3
https://gihyo.jp/list/group/Ubuntu-Weekly-Recipe
4
https://pc.watch.impress.co.jp/docs/column/ubuntu/
5
https://gihyo.jp/magazine/SD
6
https://info.nikkeibp.co.jp/media/LIN/
2
Now, a quick commercial break. Here’s a little about me.


# Page. 20

![Page Image](https://bcdn.docswell.com/page/47MYW9L27W.jpg)

Commercial Break
• Ubuntu Japanese LoCo Team member
• Embedded software engineer at SOUM
Corporation1
• Author of Ubuntu-related articles in Jpanese
• 「はじめての Ubuntu」2
• Ubuntu Weekly Recipe3
• Ubuntu 日和4
• Software Design5
• 日経 Linux6
1
https://www.soum.co.jp/
https://gihyo.jp/book/2025/978-4-297-15277-2
3
https://gihyo.jp/list/group/Ubuntu-Weekly-Recipe
4
https://pc.watch.impress.co.jp/docs/column/ubuntu/
5
https://gihyo.jp/magazine/SD
6
https://info.nikkeibp.co.jp/media/LIN/
2
I’m a member of the Ubuntu Japanese LoCo Team.


# Page. 21

![Page Image](https://bcdn.docswell.com/page/P7R9YGK5E9.jpg)

Commercial Break
• Ubuntu Japanese LoCo Team member
• Embedded software engineer at SOUM
Corporation1
• Author of Ubuntu-related articles in Jpanese
• 「はじめての Ubuntu」2
• Ubuntu Weekly Recipe3
• Ubuntu 日和4
• Software Design5
• 日経 Linux6
1
https://www.soum.co.jp/
https://gihyo.jp/book/2025/978-4-297-15277-2
3
https://gihyo.jp/list/group/Ubuntu-Weekly-Recipe
4
https://pc.watch.impress.co.jp/docs/column/ubuntu/
5
https://gihyo.jp/magazine/SD
6
https://info.nikkeibp.co.jp/media/LIN/
2
I write Ubuntu books and articles in Japanese.


# Page. 22

![Page Image](https://bcdn.docswell.com/page/PJXQGXLX7X.jpg)

My First Idea
Add subtitles to my presentation slides.
My first idea was simple: What if I added subtitles to my presentation?


# Page. 23

![Page Image](https://bcdn.docswell.com/page/3JK94WL9JD.jpg)

My First Idea
Add subtitles to my presentation slides.
• Let the subtitles compensate for my pronunciation.
• Turn the subtitles into my teleprompter.
• Let AI worry about my grammar.
Like the subtitles you can see at the bottom of the screen. They could make up for my pronunciation.


# Page. 24

![Page Image](https://bcdn.docswell.com/page/LE3WL139E5.jpg)

My First Idea
Add subtitles to my presentation slides.
• Let the subtitles compensate for my pronunciation.
• Turn the subtitles into my teleprompter.
• Let AI worry about my grammar.
They could also work as my teleprompter. I would only need to read them.


# Page. 25

![Page Image](https://bcdn.docswell.com/page/8EDKNX487G.jpg)

My First Idea
Add subtitles to my presentation slides.
• Let the subtitles compensate for my pronunciation.
• Turn the subtitles into my teleprompter.
• Let AI worry about my grammar.
AI could check my grammar beforehand.


# Page. 26

![Page Image](https://bcdn.docswell.com/page/V7PK5PMWJ8.jpg)

My First Idea
Add subtitles to my presentation slides.
• Let the subtitles compensate for my pronunciation.
• Turn the subtitles into my teleprompter.
• Let AI worry about my grammar.
But preparing subtitles takes time.


# Page. 27

![Page Image](https://bcdn.docswell.com/page/2JVVP299JQ.jpg)

Then I thought...
Then I thought


# Page. 28

![Page Image](https://bcdn.docswell.com/page/5EGLPRZ4JL.jpg)

Then I thought...
What if every presenter could have subtitles...
What if every presenter could have subtitles...


# Page. 29

![Page Image](https://bcdn.docswell.com/page/4JQY2VLR7P.jpg)

Then I thought...
What if every presenter could have subtitles...
automatically?
automatically?


# Page. 30

![Page Image](https://bcdn.docswell.com/page/K74WLMDQE1.jpg)

The Goal
“Make Real-Time Subtitles effortless for every presenter.”
So this became my goal.


# Page. 31

![Page Image](https://bcdn.docswell.com/page/LJ1YL8ZLEG.jpg)

The Goal
“Make Real-Time Subtitles effortless for every presenter.”
Make Real-Time Subtitles effortless for every presenter.


# Page. 32

![Page Image](https://bcdn.docswell.com/page/GJWGVZ9Q72.jpg)

The Goal
“Make Real-Time Subtitles effortless for every presenter.”
The key word is “effortless.” I want anyone to be able to add subtitles easily.


# Page. 33

![Page Image](https://bcdn.docswell.com/page/4EZL419373.jpg)

The Goal
“Make Real-Time Subtitles effortless for every presenter.”
The key word is “effortless.” I want anyone to be able to add subtitles easily.


# Page. 34

![Page Image](https://bcdn.docswell.com/page/Y76W3LKZ7V.jpg)

Why “local” subtitles?
The title of this talk is “Local Real-Time Subtitles.”


# Page. 35

![Page Image](https://bcdn.docswell.com/page/G75ML1P974.jpg)

Why “local” subtitles?
So why does it have to be local?


# Page. 36

![Page Image](https://bcdn.docswell.com/page/9J29L165ER.jpg)

Why “local” subtitles?
Let’s look at it from both the speaker’s and the audience’s point of view.


# Page. 37

![Page Image](https://bcdn.docswell.com/page/DEY4QZ96JM.jpg)

As a speaker
• Writing slides in English is manageable.
• Speaking beyond the slides is much harder.
• I’m never sure whether my pronunciation is clear enough.
• I don’t have time to prepare subtitles before every presentation.
As a speaker, writing slides in English isn’t too difficult.


# Page. 38

![Page Image](https://bcdn.docswell.com/page/VJNYK3D178.jpg)

As a speaker
• Writing slides in English is manageable.
• Speaking beyond the slides is much harder.
• I’m never sure whether my pronunciation is clear enough.
• I don’t have time to prepare subtitles before every presentation.
But slides can’t explain everything.


# Page. 39

![Page Image](https://bcdn.docswell.com/page/YE9PV9GYJ3.jpg)

As a speaker
• Writing slides in English is manageable.
• Speaking beyond the slides is much harder.
• I’m never sure whether my pronunciation is clear enough.
• I don’t have time to prepare subtitles before every presentation.
I also worry about whether people can understand my pronunciation.


# Page. 40

![Page Image](https://bcdn.docswell.com/page/GE8DP9VKED.jpg)

As a speaker
• Writing slides in English is manageable.
• Speaking beyond the slides is much harder.
• I’m never sure whether my pronunciation is clear enough.
• I don’t have time to prepare subtitles before every presentation.
And since I’m usually still editing my slides until the last minute, I don’t have time to prepare subtitles.


# Page. 41

![Page Image](https://bcdn.docswell.com/page/LELM1W5P7R.jpg)

As an audience member
• Reading English slides is usually fine.
• Following spoken English is much harder.
• Subtitles make talks easier to understand.
• Live translation would be even better...
• ...but it’s often expensive.
As an audience member, reading English slides is much easier than listening. At least, that’s true
for me.


# Page. 42

![Page Image](https://bcdn.docswell.com/page/4JMYW932JW.jpg)

As an audience member
• Reading English slides is usually fine.
• Following spoken English is much harder.
• Subtitles make talks easier to understand.
• Live translation would be even better...
• ...but it’s often expensive.
Keeping up with spoken English is much harder.


# Page. 43

![Page Image](https://bcdn.docswell.com/page/PJR9YGQ579.jpg)

As an audience member
• Reading English slides is usually fine.
• Following spoken English is much harder.
• Subtitles make talks easier to understand.
• Live translation would be even better...
• ...but it’s often expensive.
Subtitles make talks much easier to follow.


# Page. 44

![Page Image](https://bcdn.docswell.com/page/PEXQGX5XJX.jpg)

As an audience member
• Reading English slides is usually fine.
• Following spoken English is much harder.
• Subtitles make talks easier to understand.
• Live translation would be even better...
• ...but it’s often expensive.
Live translation would be ideal.


# Page. 45

![Page Image](https://bcdn.docswell.com/page/3EK94W69ED.jpg)

As an audience member
• Reading English slides is usually fine.
• Following spoken English is much harder.
• Subtitles make talks easier to understand.
• Live translation would be even better...
• ...but it’s often expensive.
But it’s often too expensive for event organizers.


# Page. 46

![Page Image](https://bcdn.docswell.com/page/L73WL16975.jpg)

The AI living in the cloud says...
The AI living in the cloud says.


# Page. 47

![Page Image](https://bcdn.docswell.com/page/87DKNXV8JG.jpg)

The AI living in the cloud says...
• “I can transcribe speech.”
• “I can translate it too.”
• “Just give me a powerful GPU and/or a bit of money.”
I can transcribe speech.


# Page. 48

![Page Image](https://bcdn.docswell.com/page/VJPK5PVWE8.jpg)

The AI living in the cloud says...
• “I can transcribe speech.”
• “I can translate it too.”
• “Just give me a powerful GPU and/or a bit of money.”
I can translate it too.


# Page. 49

![Page Image](https://bcdn.docswell.com/page/2EVVP2Z9EQ.jpg)

The AI living in the cloud says...
• “I can transcribe speech.”
• “I can translate it too.”
• “Just give me a powerful GPU and/or a bit of money.”
Just give me a powerful GPU and/or a bit of money.


# Page. 50

![Page Image](https://bcdn.docswell.com/page/57GLPR44EL.jpg)

Then I had an idea
What if...
Then I had another idea.


# Page. 51

![Page Image](https://bcdn.docswell.com/page/4EQY2VGRJP.jpg)

Then I had an idea
What if...
“the speaker’s laptop
could generate subtitles
during the presentation?”
The speaker’s laptop could generate subtitles during the presentation?


# Page. 52

![Page Image](https://bcdn.docswell.com/page/KJ4WLM5Q71.jpg)

Then I had an idea
What if...
“the speaker’s laptop
could generate subtitles
during the presentation?”
If that were possible, everyone would benefit.


# Page. 53

![Page Image](https://bcdn.docswell.com/page/LE1YL8WL7G.jpg)

Project Goal
Build a subtitle system that
So let’s go back to the project goal. The subtitle system should meet these requirements.


# Page. 54

![Page Image](https://bcdn.docswell.com/page/GEWGVZ6QJ2.jpg)

Project Goal
Build a subtitle system that
• is easy to use
• works with existing PDF slides
• requires no Internet connection
• runs on ordinary laptops
• is Free/Open Source Software
• runs on Ubuntu
It is easy to use.


# Page. 55

![Page Image](https://bcdn.docswell.com/page/47ZL41Y3J3.jpg)

Project Goal
Build a subtitle system that
• is easy to use
• works with existing PDF slides
• requires no Internet connection
• runs on ordinary laptops
• is Free/Open Source Software
• runs on Ubuntu
It works with existing PDF slides.


# Page. 56

![Page Image](https://bcdn.docswell.com/page/YJ6W3LDZJV.jpg)

Project Goal
Build a subtitle system that
• is easy to use
• works with existing PDF slides
• requires no Internet connection
• runs on ordinary laptops
• is Free/Open Source Software
• runs on Ubuntu
It requires no Internet connection.


# Page. 57

![Page Image](https://bcdn.docswell.com/page/GJ5ML139J4.jpg)

Project Goal
Build a subtitle system that
• is easy to use
• works with existing PDF slides
• requires no Internet connection
• runs on ordinary laptops
• is Free/Open Source Software
• runs on Ubuntu
It runs on ordinary laptops.


# Page. 58

![Page Image](https://bcdn.docswell.com/page/LE3WL162E5.jpg)

Project Goal
Build a subtitle system that
• is easy to use
• works with existing PDF slides
• requires no Internet connection
• runs on ordinary laptops
• is Free/Open Source Software
• runs on Ubuntu
It is Free/Open Source Software.


# Page. 59

![Page Image](https://bcdn.docswell.com/page/8EDKNXV67G.jpg)

Project Goal
Build a subtitle system that
• is easy to use
• works with existing PDF slides
• requires no Internet connection
• runs on ordinary laptops
• is Free/Open Source Software
• runs on Ubuntu
It runs on Ubuntu.


# Page. 60

![Page Image](https://bcdn.docswell.com/page/V7PK5PRZJ8.jpg)

There are many ways...
Speech recognition isn’t new. There are already many great solutions.
Cloud APIs
whisper.cpp
Whisper
faster-whisper
NVIDIA Parakeet
Vosk
Every solution has trade-offs.
To make this work, we first need speech recognition.
WhisperLiveKit
Moonshine


# Page. 61

![Page Image](https://bcdn.docswell.com/page/2JVVP2RMJQ.jpg)

There are many ways...
Speech recognition isn’t new. There are already many great solutions.
Cloud APIs
whisper.cpp
Whisper
faster-whisper
NVIDIA Parakeet
Vosk
Every solution has trade-offs.
Speech recognition isn’t new. There are already many great solutions.
WhisperLiveKit
Moonshine


# Page. 62

![Page Image](https://bcdn.docswell.com/page/5EGLPRMXJL.jpg)

There are many ways...
Speech recognition isn’t new. There are already many great solutions.
Whisper
Cloud APIs
whisper.cpp
faster-whisper
NVIDIA Parakeet
Vosk
Every solution has trade-offs.
But every solution has trade-offs.
WhisperLiveKit
Moonshine


# Page. 63

![Page Image](https://bcdn.docswell.com/page/4JQY2VR57P.jpg)

Moonshine, I Choose You!
Moonshine1 is designed for fast, local and real-time speech recognition.
• Fast enough for live subtitles
• Runs well on CPUs
• Also supports NPUs when available
• Easy to integrate into Python applications
• Both of the code and the English model are MIT licensed
1
https://github.com/moonshine-ai/moonshine
For this project, I chose Moonshine.


# Page. 64

![Page Image](https://bcdn.docswell.com/page/K74WLM8VE1.jpg)

Moonshine, I Choose You!
Moonshine1 is designed for fast, local and real-time speech recognition.
• Fast enough for live subtitles
• Runs well on CPUs
• Also supports NPUs when available
• Easy to integrate into Python applications
• Both of the code and the English model are MIT licensed
1
https://github.com/moonshine-ai/moonshine
Moonshine is designed for fast, local and real-time speech recognition.


# Page. 65

![Page Image](https://bcdn.docswell.com/page/LJ1YL824EG.jpg)

Moonshine, I Choose You!
Moonshine1 is designed for fast, local and real-time speech recognition.
• Fast enough for live subtitles
• Runs well on CPUs
• Also supports NPUs when available
• Easy to integrate into Python applications
• Both of the code and the English model are MIT licensed
1
https://github.com/moonshine-ai/moonshine
The biggest reason was that it works well on ordinary laptops, even without a powerful GPU or
NPU.


# Page. 66

![Page Image](https://bcdn.docswell.com/page/GJWGVZRZ72.jpg)

Why Pympress?
Pympress1 is PDF presentation software designed for dual-screen setups.
• Presenter View
• PDF-based workflow
• Easy to extend
• GPL-2.0
Instead of replacing Pympress, I extended it.
1
https://github.com/Cimbali/pympress
Next, I needed presentation software.


# Page. 67

![Page Image](https://bcdn.docswell.com/page/4EZL41RL73.jpg)

Why Pympress?
Pympress1 is PDF presentation software designed for dual-screen setups.
• Presenter View
• PDF-based workflow
• Easy to extend
• GPL-2.0
Instead of replacing Pympress, I extended it.
1
https://github.com/Cimbali/pympress
I’m using Pympress right now.


# Page. 68

![Page Image](https://bcdn.docswell.com/page/Y76W3LRM7V.jpg)

Why Pympress?
Pympress1 is PDF presentation software designed for dual-screen setups.
• Presenter View
• PDF-based workflow
• Easy to extend
• GPL-2.0
Instead of replacing Pympress, I extended it.
1
https://github.com/Cimbali/pympress
So instead of creating a new presentation tool, I decided to extend Pympress with subtitle support.


# Page. 69

![Page Image](https://bcdn.docswell.com/page/G75ML18Q74.jpg)

The Architecture
Microphone
Audio stream
Moonshine
Speech Recognition
Unix
domain Recognized text
socket
Projector
Slides + subtitles
Pympress
Subtitle Overlay
Audio is converted into subtitles and rendered directly onto the slides.
Here’s the overall architecture.


# Page. 70

![Page Image](https://bcdn.docswell.com/page/9J29L1KWER.jpg)

The Architecture
Microphone
Audio stream
Moonshine
Speech Recognition
Unix
domain Recognized text
socket
Projector
Slides + subtitles
Pympress
Subtitle Overlay
Audio is converted into subtitles and rendered directly onto the slides.
First, audio is captured from the microphone and sent to Moonshine.


# Page. 71

![Page Image](https://bcdn.docswell.com/page/DEY4QZN9JM.jpg)

The Architecture
Microphone
Audio stream
Moonshine
Speech Recognition
Unix
domain Recognized text
socket
Projector
Slides + subtitles
Pympress
Subtitle Overlay
Audio is converted into subtitles and rendered directly onto the slides.
Moonshine performs speech recognition and sends the recognized text to Pympress through a Unix
domain socket.


# Page. 72

![Page Image](https://bcdn.docswell.com/page/VJNYK3XD78.jpg)

The Architecture
Microphone
Audio stream
Moonshine
Speech Recognition
Unix
domain Recognized text
socket
Projector
Slides + subtitles
Pympress
Subtitle Overlay
Audio is converted into subtitles and rendered directly onto the slides.
Pympress displays the text as a subtitle overlay.


# Page. 73

![Page Image](https://bcdn.docswell.com/page/YE9PV988J3.jpg)

The Architecture
Microphone
Moonshine
Audio stream
Speech Recognition
Unix
domain Recognized text
socket
Projector
Pympress
Slides + subtitles
Subtitle Overlay
Audio is converted into subtitles and rendered directly onto the slides.
Finally, the slides and subtitles are shown on the projector.


# Page. 74

![Page Image](https://bcdn.docswell.com/page/GE8DP98ZED.jpg)

The Architecture
Microphone
Audio stream
Moonshine
Speech Recognition
Unix
domain Recognized text
socket
Projector
Slides + subtitles
Pympress
Subtitle Overlay
Audio is converted into subtitles and rendered directly onto the slides.
The data sent through the socket is simply JSON with text and metadata.


# Page. 75

![Page Image](https://bcdn.docswell.com/page/LELM1WR17R.jpg)

Live Demo
Now it’s time for a live demo.


# Page. 76

![Page Image](https://bcdn.docswell.com/page/4JMYW9R5JW.jpg)

Live Demo
Actually, the automatic subtitles have already been running.


# Page. 77

![Page Image](https://bcdn.docswell.com/page/PJR9YGRZ79.jpg)

Live Demo
I just kept them hidden.


# Page. 78

![Page Image](https://bcdn.docswell.com/page/PEXQGXR1JX.jpg)

Live Demo
Now I’ll turn them on.


# Page. 79

![Page Image](https://bcdn.docswell.com/page/3EK94WRMED.jpg)

Live Demo
As you can see, my speech appears on the screen with a short delay.


# Page. 80

![Page Image](https://bcdn.docswell.com/page/L73WL18275.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
This is the environment I’m using for the demo.


# Page. 81

![Page Image](https://bcdn.docswell.com/page/87DKNXM6JG.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
It’s a fairly ordinary laptop that I bought about three years ago.


# Page. 82

![Page Image](https://bcdn.docswell.com/page/VJPK5PQZE8.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
The important point is that it doesn’t have a high-end GPU.


# Page. 83

![Page Image](https://bcdn.docswell.com/page/2EVVP25MEQ.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
It also has no NPU or TPU.


# Page. 84

![Page Image](https://bcdn.docswell.com/page/57GLPR6XEL.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
And of course, it doesn’t use any cloud APIs.


# Page. 85

![Page Image](https://bcdn.docswell.com/page/4EQY2V45JP.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
Even so, it can generate subtitles at this speed with this level of accuracy.


# Page. 86

![Page Image](https://bcdn.docswell.com/page/KJ4WLMQV71.jpg)

Demo Environment
This demo is running on:
• Intel Core i5-1335U (P2 + E8, 12 Threads)
• Intel Iris Xe Graphics (iGPU)
• 16 GiB LPDDR5 RAM
• RODE VideoMicro II (3.5mm jack)
• No NPU
• No TPU
• No Internet connection
Even so, it can generate subtitles at this speed with this level of accuracy.


# Page. 87

![Page Image](https://bcdn.docswell.com/page/LE1YL8P47G.jpg)

How to Use and Configuration
Now let me briefly explain how to use it and how to configure it.


# Page. 88

![Page Image](https://bcdn.docswell.com/page/GEWGVZ2ZJ2.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
Gray text means what is speech is still being recognized.


# Page. 89

![Page Image](https://bcdn.docswell.com/page/47ZL41DLJ3.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
The recognition result is updated every 0.25 seconds.


# Page. 90

![Page Image](https://bcdn.docswell.com/page/YJ6W3L1MJV.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
White text means what is recognition is finalized.


# Page. 91

![Page Image](https://bcdn.docswell.com/page/GJ5ML14QJ4.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
Finalized subtitles stay on the screen for 1.5 seconds.


# Page. 92

![Page Image](https://bcdn.docswell.com/page/9E29L1VW7R.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
You can change this in the configuration file.


# Page. 93

![Page Image](https://bcdn.docswell.com/page/D7Y4QZ39EM.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
Speech recognition continues in the background during that time.


# Page. 94

![Page Image](https://bcdn.docswell.com/page/VENYK3VDJ8.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
Press “C” to toggle subtitles on and off.


# Page. 95

![Page Image](https://bcdn.docswell.com/page/Y79PV968E3.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
Normally, only the latest sentence is shown.


# Page. 96

![Page Image](https://bcdn.docswell.com/page/G78DP9ZZ7D.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
Press “Shift+C” to display multiple recent sentences.


# Page. 97

![Page Image](https://bcdn.docswell.com/page/L7LM1WD1JR.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
The background color, text color, position, and size can all be customized in the configuration file.


# Page. 98

![Page Image](https://bcdn.docswell.com/page/4EMYW9P5EW.jpg)

How to Use and Configuration
• Gray text: speech is still being recognized
• White text: recognition is finalized
• Press “C” to toggle subtitles on and off
• Press “Shift+C” to display up to five sentences
• Customize the layout in the configuration file (~/.config/pympress)
It looks like this.


# Page. 99

![Page Image](https://bcdn.docswell.com/page/PER9YG1ZJ9.jpg)

Installation
I’ve also documented how to install it, so you can try it yourself. I tested everything on Ubuntu
26.04 LTS.


# Page. 100

![Page Image](https://bcdn.docswell.com/page/P7XQGXY1EX.jpg)

1. Install uv and the required packages
$ curl -LsSf https://astral.sh/uv/install.sh | sh
$ sudo apt -U install pympress libportaudio2 portaudio19-dev
First, install uv and the packages needed for Pympress and Moonshine.


# Page. 101

![Page Image](https://bcdn.docswell.com/page/37K94W3M7D.jpg)

2. Install the modified version of Pympress
$ mkdir ~/Pympress &amp;&amp; cd $_
$ git clone git@github.com:m-shibata/pympress.git
$ cd pympress
$ git switch features/subtitle
Next, install my modified version of Pympress. Please use the “features/subtitle” branch.


# Page. 102

![Page Image](https://bcdn.docswell.com/page/LJ3WL1G2J5.jpg)

3. Create the Python environment with uv
$ uv venv --python /usr/bin/python3 --system-site-packages
$ PYTHONPATH=&quot;$PWD&quot; uv run pympress &lt;PDF file&gt;
Create a Python virtual environment with uv. Then launch Pympress with any PDF file. After confirming that it works, close Pympress before the next step.


# Page. 103

![Page Image](https://bcdn.docswell.com/page/8JDKNX16EG.jpg)

4. Install Moonshine
$ uv venv --python 3.12 .venv-moonshine
$ uv pip install --python .venv-moonshine/bin/python \
moonshine-voice
$ .venv-moonshine/bin/moonshine-voice download --stt --language en
$ .venv-moonshine/bin/moonshine-voice mic --language en
Next, install Moonshine. The integration script is included on the Pympress side, so you can use
the original Moonshine package.


# Page. 104

![Page Image](https://bcdn.docswell.com/page/VEPK5P5Z78.jpg)

4. Install Moonshine
$ uv venv --python 3.12 .venv-moonshine
$ uv pip install --python .venv-moonshine/bin/python \
moonshine-voice
$ .venv-moonshine/bin/moonshine-voice download --stt --language en
$ .venv-moonshine/bin/moonshine-voice mic --language en
Download the English model, then test speech recognition with the mic command. If everything
works, press Ctrl+C to stop Moonshine.


# Page. 105

![Page Image](https://bcdn.docswell.com/page/27VVP2PM7Q.jpg)

5. Launch Pympress
$ CAPTION_SOCKET=\
&quot;${XDG_RUNTIME_DIR:-/tmp}/pympress-caption-$UID.sock&quot;
$ PYTHONPATH=&quot;$PWD&quot; uv run pympress \
--caption-socket &quot;$CAPTION_SOCKET&quot; &lt;PDF file&gt;
Now launch the subtitle-enabled version of Pympress. Don’t forget to specify the Unix domain
socket with the “--caption-socket” option.


# Page. 106

![Page Image](https://bcdn.docswell.com/page/5JGLPRPX7L.jpg)

6. Launch Moonshine
$ CAPTION_SOCKET=\
&quot;${XDG_RUNTIME_DIR:-/tmp}/pympress-caption-$UID.sock&quot;
$ PYTHONPATH=&quot;$PWD&quot; \
.venv-moonshine/bin/python scripts/moonshine_captions.py \
--socket &quot;$CAPTION_SOCKET&quot; --language en \
--update-interval 0.25
Next, open another terminal and launch Moonshine. Specify the same Unix domain socket with
the “--socket” option.


# Page. 107

![Page Image](https://bcdn.docswell.com/page/47QY2V25EP.jpg)

6. Launch Moonshine
$ CAPTION_SOCKET=\
&quot;${XDG_RUNTIME_DIR:-/tmp}/pympress-caption-$UID.sock&quot;
$ PYTHONPATH=&quot;$PWD&quot; \
.venv-moonshine/bin/python scripts/moonshine_captions.py \
--socket &quot;$CAPTION_SOCKET&quot; --language en \
--update-interval 0.25
The “--update-interval” option controls how often the recognition result is updated. A shorter interval makes the subtitles more responsive, but the intermediate results may be less accurate.


# Page. 108

![Page Image](https://bcdn.docswell.com/page/KE4WLMLVJ1.jpg)

Implementation Details
Need more details?
You might be wondering how it’s implemented. Would you like a deeper explanation?


# Page. 109

![Page Image](https://bcdn.docswell.com/page/L71YL8L4JG.jpg)

Implementation Details
Need more details?
Ask your favorite Coding Agent.
Just ask your favorite Coding Agent.


# Page. 110

![Page Image](https://bcdn.docswell.com/page/G7WGVZVZE2.jpg)

Implementation Details
Need more details?
Ask your favorite Coding Agent.
https://github.com/m-shibata/pympress
Actually, the implementation is quite simple.


# Page. 111

![Page Image](https://bcdn.docswell.com/page/4JZL414LE3.jpg)

Implementation Details
Need more details?
Ask your favorite Coding Agent.
https://github.com/m-shibata/pympress
Since Pympress is written in Python with GTK, I simply added an overlay widget using Gtk.Overlay.


# Page. 112

![Page Image](https://bcdn.docswell.com/page/YE6W3L3MEV.jpg)

What About Q&amp;A?
Real-time sbutitles can help me speak.
But I still have to understand your questions...
Since Pympress is written in Python with GTK, I simply added an overlay widget using Gtk.Overlay.


# Page. 113

![Page Image](https://bcdn.docswell.com/page/GE5ML1LQE4.jpg)

What About Q&amp;A?
Real-time sbutitles can help me speak.
But I still have to understand your questions...
There’s a proble for future me.
Since Pympress is written in Python with GTK, I simply added an overlay widget using Gtk.Overlay.


# Page. 114

![Page Image](https://bcdn.docswell.com/page/9729L1LWJR.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
Here are some ideas for future improvements.


# Page. 115

![Page Image](https://bcdn.docswell.com/page/DJY4QZQ97M.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
I’d love to support Japanese speech recognition as well.


# Page. 116

![Page Image](https://bcdn.docswell.com/page/V7NYK3KDE8.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
Moonshine already provides Japanese model, but unlike the English model, they aren’t released
under the MIT license.


# Page. 117

![Page Image](https://bcdn.docswell.com/page/YJ9PV9V873.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
Once Japanese recognition works well, automatic translation would be the next step. I’m also interested in finding a translation engine that runs efficiently on CPUs.


# Page. 118

![Page Image](https://bcdn.docswell.com/page/GJ8DP9PZJD.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
If we can also generate translated speech, we’ll have a complete speech translation system.


# Page. 119

![Page Image](https://bcdn.docswell.com/page/LJLM1W11ER.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
I’d also like to support more speech recognition engines. The communication protocol with Pympress is simple, so adding new engines shouldn’t be difficult.


# Page. 120

![Page Image](https://bcdn.docswell.com/page/47MYW9W57W.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
And if everything goes well, I’d like to propose these changes upstream.


# Page. 121

![Page Image](https://bcdn.docswell.com/page/P7R9YGYZE9.jpg)

What’s Next?
• Better Japanese speech recognition
Moonshine already provides Japanese model, but under a non-MIT license.
• Automatic subtitle translation
Using Marian or Argos Translate?
• Automatic speech translation
Moonshine provides TTS capabilities, or use Piper
• Support for additional speech recognition engines
But all of this development is presentation-driven. Unless I submit another conference proposal,
none of it will happen.


# Page. 122

![Page Image](https://bcdn.docswell.com/page/PJXQGXG17X.jpg)

Conclusion
Finally, let me wrap up.


# Page. 123

![Page Image](https://bcdn.docswell.com/page/3JK94W4MJD.jpg)

Conclusion
Accessibility should not depend on cloud services.


