OMERUTA
Docs

Dictate text and record voice notes

Omeruta provides two voice features in editable Markdown notes.

Omeruta provides two voice features in editable Markdown notes:

  • Dictate converts your speech into text at the current cursor position.
  • Voice note records your speech as audio, saves it, and inserts a playable audio block into the note.

Voice notes and dictation are separate. Starting a voice note does not insert a live transcript, and ordinary dictation does not save its audio unless you turn on Save dictation audio.

Set up dictation

Voice notes can be recorded without installing a speech model. Dictation and audio-block transcription use a local Whisper speech model.

  1. Open Settings, then select AI.
  2. Expand Dictation & captions.
  3. Turn on Voice dictation.
  4. Choose a Model. If it is not installed, select Download model and wait for the status to show Installed.
  5. Choose a Mic. Leave it on System default to use the microphone selected by your operating system.
  6. If the selected model supports multiple languages, choose a Language or Auto-detect.

Omeruta performs the speech-to-text step on this device. The first voice session can trigger an operating-system microphone request; microphone access must be allowed before recording can begin.

Dictate text into a note

  1. Open an editable Markdown note.
  2. Place the cursor where the dictated text should begin.
  3. Select Dictate in the document action bar.
  4. Speak normally. When Live transcript while dictating is on, provisional words appear at the cursor while you speak.
  5. Select Pause dictation when you need a break. Select Resume dictation to continue.
  6. Select Stop dictation. Omeruta finishes the queued speech and keeps the final text in the note.

The text is ordinary note content after it is inserted, so you can edit, format, move, or delete it like text you typed.

The discard control stops the session immediately. Text already inserted into the note stays. If Save dictation audio is on, the unfinished audio is discarded instead of being saved.

Save dictation audio

Save dictation audio is off by default. Turn it on under Settings → AI → Dictation & captions when you want Omeruta to keep the audio source as well as the inserted text.

After you stop dictation, the audio appears under Media Library → Voice. Saved dictation audio is linked to its source note, but it is not inserted as an audio block inside that note. Dictation performed while audio saving is off does not create an item in the Voice section.

Record a voice note

  1. Open an editable Markdown note and place the cursor where the audio block should appear.
  2. Select Record a voice note in the document action bar.
  3. Speak. Use Pause voice note and Resume voice note when needed.
  4. Select Stop and save voice note.

Omeruta saves the recording in the active workspace's Media Library and inserts a playable audio block at the original cursor position. Voice notes are always saved, regardless of the Save dictation audio setting.

Select Cancel voice note and discard audio if you do not want to keep the recording. Cancelling does not insert an audio block.

Transcribe a saved voice note

The audio block created by a voice note has its own transcription controls:

  1. Choose the spoken language from Transcription language.
  2. Select Transcribe audio.
  3. Wait for transcription to finish.

Omeruta inserts the transcript below the audio block in the same note. The inserted result contains readable transcript text and a timestamped list. It is editable note content.

Important behavior

  • Voice features are unavailable in read-only pages, including the pages in Tutorial & documentation.
  • Keep the target note open until the voice session finishes. Changing the active file or workspace stops the session and discards unfinished capture.
  • Only one Omeruta voice session can use the microphone at a time. Omeruta also reports when another application is using the selected microphone.
  • A retained voice session stops and saves automatically if it reaches four hours.