Local Real-Time Subtitles on Ubuntu with Moonshine and Pympress

>100 Views

August 09, 26

スライド概要

Language barriers remain a significant challenge in international conferences. Not all speakers are fluent in English, and even when they are, their speech may not always be easy for diverse audiences to understand. While some events provide professional captioning or simultaneous interpretation, such services are not always available for every session.

In this talk, I will present a self-contained approach that enables speakers to generate real-time subtitles directly from their own laptop. By combining Moonshine Voice, a lightweight and efficient speech-to-text (STT) system that can run even on devices like Raspberry Pi, with Pympress, a PDF-based presentation tool, we can build a fully local, automatic subtitle system on Ubuntu.

This approach allows speakers to make their talks more accessible without relying on external infrastructure or services, making it easier to participate in international events regardless of location or available resources.

https://events.canonical.com/event/146/contributions/935/

profile-image

Ubuntu Japanese LoCo Team member

シェア

またはPlayer版

埋め込む »CMSなどでJSが使えない場合

ダウンロード

関連スライド

各ページのテキスト
1.

Local Real-Time Subtitles on Ubuntu with Moonshine and Pympress Mitsuya Shibata Ubuntu Japanese Team / SOUM Corporation 2026-08-09

2.

Introduction

3.

Two Questions • Can you pronounce English well? • Can you easily understand spoken English? First, I’d like to ask you two questions.

4.

Two Questions • Can you pronounce English well? • Can you easily understand spoken English? If you’re attending UbuCon Asia and COSCUP, you can probably communicate in English.

5.

Two Questions • Can you pronounce English well? • Can you easily understand spoken English? But I can’t! If you’re attending UbuCon Asia and COSCUP, you can probably communicate in English.

6.

Today’s topics That’s all.

7.

Today’s topics That’s all. Thank you very much!

8.

The English Cliff International Tech Community Summit (UbuCon Asia, COSCUP) (・∀・) 人 (≧▽≦) ------------------------+ | | | The English Cliff | | | | (´・ω ・`) < "I can't speak English..." +------------------------------------------At international technical communities like UbuCon Asia, English is the common language. There is a language barrier.

9.

The English Cliff International Tech Community Summit (UbuCon Asia, COSCUP) (・∀・) 人 (≧▽≦) ------------------------+ | | | The English Cliff | | | | (´・ω ・`) < "I can't speak English..." +------------------------------------------If you can’t speak or understand English, it’s hard to communicate in person.

10.
[beta]
The English Cliff
International Tech Community Summit
(UbuCon Asia, COSCUP)
(・∀・) 人 (≧▽≦)
------------------------+
|
|
|
The English Cliff |
|
|
| (´・ω ・`) < "I can't speak English..."
+------------------------------------------So I started wondering: Could open source help us overcome this barrier?

11.

Actually... It’s not just English. To be honest, it isn’t only about English.

12.

Actually... It’s not just English. I’m not even very good at communicating in my native language. I’m not very good at communicating even in my native language.

13.

Actually... It’s not just English. I’m not even very good at communicating in my native language. I’m not even good at saying “hello”.

14.

The Real Problem International Tech Community Summit (UbuCon Asia, COSCUP) (・∀・) 人 (≧▽≦) ------------------------+ | The English Cliff | | (´・ω ・`) < "I can't speak English..." +---------------------+ | The Confidence Cliff | | Me < "H-h-hello..." +--------------------So this is the real problem.

15.

The Real Problem International Tech Community Summit (UbuCon Asia, COSCUP) (・∀・) 人 (≧▽≦) ------------------------+ | The English Cliff | | (´・ω ・`) < "I can't speak English..." +---------------------+ | The Confidence Cliff | | Me < "H-h-hello..." +--------------------I have to overcome not only the English Cliff, but also the Confidence Cliff.

16.

But...

17.

But... I still wanted to speak at UbuCon Asia.

18.

But... I still wanted to speak at UbuCon Asia. That’s how this project began. I wanted to find a way to give a talk in English.

19.

Commercial Break • Ubuntu Japanese LoCo Team member • Embedded software engineer at SOUM Corporation1 • Author of Ubuntu-related articles in Jpanese • 「はじめての Ubuntu」2 • Ubuntu Weekly Recipe3 • Ubuntu 日和4 • Software Design5 • 日経 Linux6 1 https://www.soum.co.jp/ https://gihyo.jp/book/2025/978-4-297-15277-2 3 https://gihyo.jp/list/group/Ubuntu-Weekly-Recipe 4 https://pc.watch.impress.co.jp/docs/column/ubuntu/ 5 https://gihyo.jp/magazine/SD 6 https://info.nikkeibp.co.jp/media/LIN/ 2 Now, a quick commercial break. Here’s a little about me.

20.

Commercial Break • Ubuntu Japanese LoCo Team member • Embedded software engineer at SOUM Corporation1 • Author of Ubuntu-related articles in Jpanese • 「はじめての Ubuntu」2 • Ubuntu Weekly Recipe3 • Ubuntu 日和4 • Software Design5 • 日経 Linux6 1 https://www.soum.co.jp/ https://gihyo.jp/book/2025/978-4-297-15277-2 3 https://gihyo.jp/list/group/Ubuntu-Weekly-Recipe 4 https://pc.watch.impress.co.jp/docs/column/ubuntu/ 5 https://gihyo.jp/magazine/SD 6 https://info.nikkeibp.co.jp/media/LIN/ 2 I’m a member of the Ubuntu Japanese LoCo Team.

21.

Commercial Break • Ubuntu Japanese LoCo Team member • Embedded software engineer at SOUM Corporation1 • Author of Ubuntu-related articles in Jpanese • 「はじめての Ubuntu」2 • Ubuntu Weekly Recipe3 • Ubuntu 日和4 • Software Design5 • 日経 Linux6 1 https://www.soum.co.jp/ https://gihyo.jp/book/2025/978-4-297-15277-2 3 https://gihyo.jp/list/group/Ubuntu-Weekly-Recipe 4 https://pc.watch.impress.co.jp/docs/column/ubuntu/ 5 https://gihyo.jp/magazine/SD 6 https://info.nikkeibp.co.jp/media/LIN/ 2 I write Ubuntu books and articles in Japanese.

22.

My First Idea Add subtitles to my presentation slides. My first idea was simple: What if I added subtitles to my presentation?

23.

My First Idea Add subtitles to my presentation slides. • Let the subtitles compensate for my pronunciation. • Turn the subtitles into my teleprompter. • Let AI worry about my grammar. Like the subtitles you can see at the bottom of the screen. They could make up for my pronunciation.

24.

My First Idea Add subtitles to my presentation slides. • Let the subtitles compensate for my pronunciation. • Turn the subtitles into my teleprompter. • Let AI worry about my grammar. They could also work as my teleprompter. I would only need to read them.

25.

My First Idea Add subtitles to my presentation slides. • Let the subtitles compensate for my pronunciation. • Turn the subtitles into my teleprompter. • Let AI worry about my grammar. AI could check my grammar beforehand.

26.

My First Idea Add subtitles to my presentation slides. • Let the subtitles compensate for my pronunciation. • Turn the subtitles into my teleprompter. • Let AI worry about my grammar. But preparing subtitles takes time.

27.

Then I thought... Then I thought

28.

Then I thought... What if every presenter could have subtitles... What if every presenter could have subtitles...

29.

Then I thought... What if every presenter could have subtitles... automatically? automatically?

30.

The Goal “Make Real-Time Subtitles effortless for every presenter.” So this became my goal.

31.

The Goal “Make Real-Time Subtitles effortless for every presenter.” Make Real-Time Subtitles effortless for every presenter.

32.

The Goal “Make Real-Time Subtitles effortless for every presenter.” The key word is “effortless.” I want anyone to be able to add subtitles easily.

33.

The Goal “Make Real-Time Subtitles effortless for every presenter.” The key word is “effortless.” I want anyone to be able to add subtitles easily.

34.

Why “local” subtitles? The title of this talk is “Local Real-Time Subtitles.”

35.

Why “local” subtitles? So why does it have to be local?

36.

Why “local” subtitles? Let’s look at it from both the speaker’s and the audience’s point of view.

37.

As a speaker • Writing slides in English is manageable. • Speaking beyond the slides is much harder. • I’m never sure whether my pronunciation is clear enough. • I don’t have time to prepare subtitles before every presentation. As a speaker, writing slides in English isn’t too difficult.

38.

As a speaker • Writing slides in English is manageable. • Speaking beyond the slides is much harder. • I’m never sure whether my pronunciation is clear enough. • I don’t have time to prepare subtitles before every presentation. But slides can’t explain everything.

39.

As a speaker • Writing slides in English is manageable. • Speaking beyond the slides is much harder. • I’m never sure whether my pronunciation is clear enough. • I don’t have time to prepare subtitles before every presentation. I also worry about whether people can understand my pronunciation.

40.

As a speaker • Writing slides in English is manageable. • Speaking beyond the slides is much harder. • I’m never sure whether my pronunciation is clear enough. • I don’t have time to prepare subtitles before every presentation. And since I’m usually still editing my slides until the last minute, I don’t have time to prepare subtitles.

41.

As an audience member • Reading English slides is usually fine. • Following spoken English is much harder. • Subtitles make talks easier to understand. • Live translation would be even better... • ...but it’s often expensive. As an audience member, reading English slides is much easier than listening. At least, that’s true for me.

42.

As an audience member • Reading English slides is usually fine. • Following spoken English is much harder. • Subtitles make talks easier to understand. • Live translation would be even better... • ...but it’s often expensive. Keeping up with spoken English is much harder.

43.

As an audience member • Reading English slides is usually fine. • Following spoken English is much harder. • Subtitles make talks easier to understand. • Live translation would be even better... • ...but it’s often expensive. Subtitles make talks much easier to follow.

44.

As an audience member • Reading English slides is usually fine. • Following spoken English is much harder. • Subtitles make talks easier to understand. • Live translation would be even better... • ...but it’s often expensive. Live translation would be ideal.

45.

As an audience member • Reading English slides is usually fine. • Following spoken English is much harder. • Subtitles make talks easier to understand. • Live translation would be even better... • ...but it’s often expensive. But it’s often too expensive for event organizers.

46.

The AI living in the cloud says... The AI living in the cloud says.

47.

The AI living in the cloud says... • “I can transcribe speech.” • “I can translate it too.” • “Just give me a powerful GPU and/or a bit of money.” I can transcribe speech.

48.

The AI living in the cloud says... • “I can transcribe speech.” • “I can translate it too.” • “Just give me a powerful GPU and/or a bit of money.” I can translate it too.

49.

The AI living in the cloud says... • “I can transcribe speech.” • “I can translate it too.” • “Just give me a powerful GPU and/or a bit of money.” Just give me a powerful GPU and/or a bit of money.

50.

Then I had an idea What if... Then I had another idea.

51.

Then I had an idea What if... “the speaker’s laptop could generate subtitles during the presentation?” The speaker’s laptop could generate subtitles during the presentation?

52.

Then I had an idea What if... “the speaker’s laptop could generate subtitles during the presentation?” If that were possible, everyone would benefit.

53.

Project Goal Build a subtitle system that So let’s go back to the project goal. The subtitle system should meet these requirements.

54.

Project Goal Build a subtitle system that • is easy to use • works with existing PDF slides • requires no Internet connection • runs on ordinary laptops • is Free/Open Source Software • runs on Ubuntu It is easy to use.

55.

Project Goal Build a subtitle system that • is easy to use • works with existing PDF slides • requires no Internet connection • runs on ordinary laptops • is Free/Open Source Software • runs on Ubuntu It works with existing PDF slides.

56.

Project Goal Build a subtitle system that • is easy to use • works with existing PDF slides • requires no Internet connection • runs on ordinary laptops • is Free/Open Source Software • runs on Ubuntu It requires no Internet connection.

57.

Project Goal Build a subtitle system that • is easy to use • works with existing PDF slides • requires no Internet connection • runs on ordinary laptops • is Free/Open Source Software • runs on Ubuntu It runs on ordinary laptops.

58.

Project Goal Build a subtitle system that • is easy to use • works with existing PDF slides • requires no Internet connection • runs on ordinary laptops • is Free/Open Source Software • runs on Ubuntu It is Free/Open Source Software.

59.

Project Goal Build a subtitle system that • is easy to use • works with existing PDF slides • requires no Internet connection • runs on ordinary laptops • is Free/Open Source Software • runs on Ubuntu It runs on Ubuntu.

60.

There are many ways... Speech recognition isn’t new. There are already many great solutions. Cloud APIs whisper.cpp Whisper faster-whisper NVIDIA Parakeet Vosk Every solution has trade-offs. To make this work, we first need speech recognition. WhisperLiveKit Moonshine

61.

There are many ways... Speech recognition isn’t new. There are already many great solutions. Cloud APIs whisper.cpp Whisper faster-whisper NVIDIA Parakeet Vosk Every solution has trade-offs. Speech recognition isn’t new. There are already many great solutions. WhisperLiveKit Moonshine

62.

There are many ways... Speech recognition isn’t new. There are already many great solutions. Whisper Cloud APIs whisper.cpp faster-whisper NVIDIA Parakeet Vosk Every solution has trade-offs. But every solution has trade-offs. WhisperLiveKit Moonshine

63.

Moonshine, I Choose You! Moonshine1 is designed for fast, local and real-time speech recognition. • Fast enough for live subtitles • Runs well on CPUs • Also supports NPUs when available • Easy to integrate into Python applications • Both of the code and the English model are MIT licensed 1 https://github.com/moonshine-ai/moonshine For this project, I chose Moonshine.

64.

Moonshine, I Choose You! Moonshine1 is designed for fast, local and real-time speech recognition. • Fast enough for live subtitles • Runs well on CPUs • Also supports NPUs when available • Easy to integrate into Python applications • Both of the code and the English model are MIT licensed 1 https://github.com/moonshine-ai/moonshine Moonshine is designed for fast, local and real-time speech recognition.

65.

Moonshine, I Choose You! Moonshine1 is designed for fast, local and real-time speech recognition. • Fast enough for live subtitles • Runs well on CPUs • Also supports NPUs when available • Easy to integrate into Python applications • Both of the code and the English model are MIT licensed 1 https://github.com/moonshine-ai/moonshine The biggest reason was that it works well on ordinary laptops, even without a powerful GPU or NPU.

66.

Why Pympress? Pympress1 is PDF presentation software designed for dual-screen setups. • Presenter View • PDF-based workflow • Easy to extend • GPL-2.0 Instead of replacing Pympress, I extended it. 1 https://github.com/Cimbali/pympress Next, I needed presentation software.

67.

Why Pympress? Pympress1 is PDF presentation software designed for dual-screen setups. • Presenter View • PDF-based workflow • Easy to extend • GPL-2.0 Instead of replacing Pympress, I extended it. 1 https://github.com/Cimbali/pympress I’m using Pympress right now.

68.

Why Pympress? Pympress1 is PDF presentation software designed for dual-screen setups. • Presenter View • PDF-based workflow • Easy to extend • GPL-2.0 Instead of replacing Pympress, I extended it. 1 https://github.com/Cimbali/pympress So instead of creating a new presentation tool, I decided to extend Pympress with subtitle support.

69.

The Architecture Microphone Audio stream Moonshine Speech Recognition Unix domain Recognized text socket Projector Slides + subtitles Pympress Subtitle Overlay Audio is converted into subtitles and rendered directly onto the slides. Here’s the overall architecture.

70.

The Architecture Microphone Audio stream Moonshine Speech Recognition Unix domain Recognized text socket Projector Slides + subtitles Pympress Subtitle Overlay Audio is converted into subtitles and rendered directly onto the slides. First, audio is captured from the microphone and sent to Moonshine.

71.

The Architecture Microphone Audio stream Moonshine Speech Recognition Unix domain Recognized text socket Projector Slides + subtitles Pympress Subtitle Overlay Audio is converted into subtitles and rendered directly onto the slides. Moonshine performs speech recognition and sends the recognized text to Pympress through a Unix domain socket.

72.

The Architecture Microphone Audio stream Moonshine Speech Recognition Unix domain Recognized text socket Projector Slides + subtitles Pympress Subtitle Overlay Audio is converted into subtitles and rendered directly onto the slides. Pympress displays the text as a subtitle overlay.

73.

The Architecture Microphone Moonshine Audio stream Speech Recognition Unix domain Recognized text socket Projector Pympress Slides + subtitles Subtitle Overlay Audio is converted into subtitles and rendered directly onto the slides. Finally, the slides and subtitles are shown on the projector.

74.

The Architecture Microphone Audio stream Moonshine Speech Recognition Unix domain Recognized text socket Projector Slides + subtitles Pympress Subtitle Overlay Audio is converted into subtitles and rendered directly onto the slides. The data sent through the socket is simply JSON with text and metadata.

75.

Live Demo Now it’s time for a live demo.

76.

Live Demo Actually, the automatic subtitles have already been running.

77.

Live Demo I just kept them hidden.

78.

Live Demo Now I’ll turn them on.

79.

Live Demo As you can see, my speech appears on the screen with a short delay.

80.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection This is the environment I’m using for the demo.

81.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection It’s a fairly ordinary laptop that I bought about three years ago.

82.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection The important point is that it doesn’t have a high-end GPU.

83.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection It also has no NPU or TPU.

84.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection And of course, it doesn’t use any cloud APIs.

85.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection Even so, it can generate subtitles at this speed with this level of accuracy.

86.

Demo Environment This demo is running on: • Intel Core i5-1335U (P2 + E8, 12 Threads) • Intel Iris Xe Graphics (iGPU) • 16 GiB LPDDR5 RAM • RODE VideoMicro II (3.5mm jack) • No NPU • No TPU • No Internet connection Even so, it can generate subtitles at this speed with this level of accuracy.

87.

How to Use and Configuration Now let me briefly explain how to use it and how to configure it.

88.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) Gray text means what is speech is still being recognized.

89.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) The recognition result is updated every 0.25 seconds.

90.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) White text means what is recognition is finalized.

91.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) Finalized subtitles stay on the screen for 1.5 seconds.

92.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) You can change this in the configuration file.

93.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) Speech recognition continues in the background during that time.

94.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) Press “C” to toggle subtitles on and off.

95.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) Normally, only the latest sentence is shown.

96.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) Press “Shift+C” to display multiple recent sentences.

97.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) The background color, text color, position, and size can all be customized in the configuration file.

98.

How to Use and Configuration • Gray text: speech is still being recognized • White text: recognition is finalized • Press “C” to toggle subtitles on and off • Press “Shift+C” to display up to five sentences • Customize the layout in the configuration file (~/.config/pympress) It looks like this.

99.

Installation I’ve also documented how to install it, so you can try it yourself. I tested everything on Ubuntu 26.04 LTS.

100.

1. Install uv and the required packages $ curl -LsSf https://astral.sh/uv/install.sh | sh $ sudo apt -U install pympress libportaudio2 portaudio19-dev First, install uv and the packages needed for Pympress and Moonshine.

101.

2. Install the modified version of Pympress $ mkdir ~/Pympress && cd $_ $ git clone [email protected]:m-shibata/pympress.git $ cd pympress $ git switch features/subtitle Next, install my modified version of Pympress. Please use the “features/subtitle” branch.

102.

3. Create the Python environment with uv $ uv venv --python /usr/bin/python3 --system-site-packages $ PYTHONPATH="$PWD" uv run pympress <PDF file> Create a Python virtual environment with uv. Then launch Pympress with any PDF file. After confirming that it works, close Pympress before the next step.

103.

4. Install Moonshine $ uv venv --python 3.12 .venv-moonshine $ uv pip install --python .venv-moonshine/bin/python \ moonshine-voice $ .venv-moonshine/bin/moonshine-voice download --stt --language en $ .venv-moonshine/bin/moonshine-voice mic --language en Next, install Moonshine. The integration script is included on the Pympress side, so you can use the original Moonshine package.

104.

4. Install Moonshine $ uv venv --python 3.12 .venv-moonshine $ uv pip install --python .venv-moonshine/bin/python \ moonshine-voice $ .venv-moonshine/bin/moonshine-voice download --stt --language en $ .venv-moonshine/bin/moonshine-voice mic --language en Download the English model, then test speech recognition with the mic command. If everything works, press Ctrl+C to stop Moonshine.

105.
[beta]
5. Launch Pympress

$ CAPTION_SOCKET=\
"${XDG_RUNTIME_DIR:-/tmp}/pympress-caption-$UID.sock"
$ PYTHONPATH="$PWD" uv run pympress \
--caption-socket "$CAPTION_SOCKET" <PDF file>

Now launch the subtitle-enabled version of Pympress. Don’t forget to specify the Unix domain
socket with the “--caption-socket” option.

106.
[beta]
6. Launch Moonshine

$ CAPTION_SOCKET=\
"${XDG_RUNTIME_DIR:-/tmp}/pympress-caption-$UID.sock"
$ PYTHONPATH="$PWD" \
.venv-moonshine/bin/python scripts/moonshine_captions.py \
--socket "$CAPTION_SOCKET" --language en \
--update-interval 0.25

Next, open another terminal and launch Moonshine. Specify the same Unix domain socket with
the “--socket” option.

107.
[beta]
6. Launch Moonshine

$ CAPTION_SOCKET=\
"${XDG_RUNTIME_DIR:-/tmp}/pympress-caption-$UID.sock"
$ PYTHONPATH="$PWD" \
.venv-moonshine/bin/python scripts/moonshine_captions.py \
--socket "$CAPTION_SOCKET" --language en \
--update-interval 0.25

The “--update-interval” option controls how often the recognition result is updated. A shorter interval makes the subtitles more responsive, but the intermediate results may be less accurate.

108.

Implementation Details Need more details? You might be wondering how it’s implemented. Would you like a deeper explanation?

109.

Implementation Details Need more details? Ask your favorite Coding Agent. Just ask your favorite Coding Agent.

110.

Implementation Details Need more details? Ask your favorite Coding Agent. https://github.com/m-shibata/pympress Actually, the implementation is quite simple.

111.

Implementation Details Need more details? Ask your favorite Coding Agent. https://github.com/m-shibata/pympress Since Pympress is written in Python with GTK, I simply added an overlay widget using Gtk.Overlay.

112.

What About Q&A? Real-time sbutitles can help me speak. But I still have to understand your questions... Since Pympress is written in Python with GTK, I simply added an overlay widget using Gtk.Overlay.

113.

What About Q&A? Real-time sbutitles can help me speak. But I still have to understand your questions... There’s a proble for future me. Since Pympress is written in Python with GTK, I simply added an overlay widget using Gtk.Overlay.

114.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines Here are some ideas for future improvements.

115.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines I’d love to support Japanese speech recognition as well.

116.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines Moonshine already provides Japanese model, but unlike the English model, they aren’t released under the MIT license.

117.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines Once Japanese recognition works well, automatic translation would be the next step. I’m also interested in finding a translation engine that runs efficiently on CPUs.

118.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines If we can also generate translated speech, we’ll have a complete speech translation system.

119.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines I’d also like to support more speech recognition engines. The communication protocol with Pympress is simple, so adding new engines shouldn’t be difficult.

120.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines And if everything goes well, I’d like to propose these changes upstream.

121.

What’s Next? • Better Japanese speech recognition Moonshine already provides Japanese model, but under a non-MIT license. • Automatic subtitle translation Using Marian or Argos Translate? • Automatic speech translation Moonshine provides TTS capabilities, or use Piper • Support for additional speech recognition engines But all of this development is presentation-driven. Unless I submit another conference proposal, none of it will happen.

122.

Conclusion Finally, let me wrap up.

123.

Conclusion Accessibility should not depend on cloud services.