You should not need a translator to buy a tool whose whole job is turning your words into text. This page defines every term you will meet while comparing dictation software — including the ones in our own settings screens. Each definition is one sentence; the details after it are optional reading.

The basics

Dictation

Dictation is speaking out loud while software turns your words into written text. People also call it voice typing or speech-to-text. All three names mean the same thing: you talk, the computer types.

Voice typing

Voice typing is another name for dictation — talking instead of typing, with the words appearing wherever your text cursor is. Windows calls its built-in tool "voice typing" (the Win+H feature). CaringDictate does the same job in most programs, without the built-in tool's time-outs.

Speech-to-text

Speech-to-text is the technology that converts spoken audio into written words. You will see it abbreviated as STT in technical articles. When a product says "powered by speech-to-text," it means dictation.

Speech recognition

Speech recognition is the broader field of computers understanding human speech — both what you said (dictation) and what you meant (commands). A dictation app uses speech recognition to write your words down. A voice assistant uses it to act on them.

Transcription

Transcription is turning a recording of speech — a meeting, an interview, a voicemail — into a written document after the fact. Dictation is live (you watch words appear as you speak); transcription usually happens later, from a file. Tools like Otter.ai are transcription tools, not dictation tools.

How recognition works

Voice model (recognition model)

A voice model is the trained brain of a dictation program — the file that actually understands speech. Bigger models are usually more accurate but need a more powerful computer. CaringDictate picks the model size that fits your machine, from about 190 MB to about 3.1 GB.

Whisper

Whisper is an open-source family of voice models released by OpenAI in 2022, known for near-human accuracy across many languages. It is the engine behind many modern dictation and transcription products, including CaringDictate. Because it is open source, it can run entirely on your own computer.

On-device (offline) recognition

On-device recognition means your speech is processed on your own computer — the audio never travels to a company's server. This works with no internet, keeps private thoughts private, and never stops working because a service is down. It is how CaringDictate runs.

Cloud recognition

Cloud recognition means your voice is streamed over the internet to a company's servers, converted to text there, and sent back. Recognition happens on a server rather than on your machine. Where any given tool does its recognition is set out in its own documentation, and some products offer both — it is worth checking rather than assuming.

Accuracy (word error rate)

Accuracy is the percentage of words a dictation program gets right; word error rate (WER) is the same idea measured backwards — the percentage it gets wrong. Modern Whisper-based dictation is highly accurate on clear English speech. Results vary with microphone quality, background noise, and how clearly you speak.

Latency

Latency is the delay between finishing a sentence and seeing the text appear. Under about a third of a second feels instant; much past that feels laggy. On-device recognition avoids the round-trip to a server, which helps.

Automatic language detection

Automatic language detection means the software figures out which language you are speaking without being told. Useful if you switch between languages mid-day. CaringDictate's "Automatic" setting does this across its 20 supported languages.

Using a dictation app day to day

Talk button (hotkey)

The talk button is the single key or mouse button that starts and stops dictation. In CaringDictate the default is the Right Ctrl key, and you can change it to any key or mouse button that is comfortable for your hands.

Push-to-talk (hold to talk)

Push-to-talk means holding the talk button down while you speak and releasing it when you finish — like a walkie-talkie. Good for short bursts: a search box, a quick reply. The alternative is tap-to-toggle, better for longer writing.

Toggle (tap to talk)

Toggle mode means tapping the talk button once to start listening and tapping again to stop. Your hands are completely free in between — helpful when holding a key is uncomfortable, or when you dictate long passages.

Wake word (voice activation)

A wake word is a phrase you say out loud — like "start dictating" — that begins dictation without touching the keyboard or mouse at all. This matters most when reaching the keyboard is the hard part. CaringDictate added user-settable voice activation in version 3.5.

Listening indicator

The listening indicator is the small on-screen signal that shows the microphone is live and your words are being heard. A trustworthy indicator matters: you should never have to wonder whether the computer is listening.

Voice commands

Voice commands are spoken instructions — like "new line" or "delete that" — that control formatting instead of writing words. Dictation apps differ a lot here. Simple commands cover most everyday writing; heavy hands-free computer control is a separate category (see Voice Access).

Punctuation commands

Punctuation commands are saying "period," "comma," or "question mark" out loud to insert punctuation. Modern engines like Whisper also punctuate automatically from your phrasing and pauses, so you can often just talk naturally.

Custom dictionary (word list)

A custom dictionary is your personal list of names and special words the recognizer should trust — family names, medication names, industry jargon. Adding "Zofran" or "Oaxaca" once beats correcting it forever. CaringDictate includes one under Writing & words.

Text insertion

Text insertion is how dictated words physically arrive in your document — the app either types them keystroke by keystroke or pastes them in one piece. Pasting whole phrases is faster and far more reliable across different programs, which is why CaringDictate delivers text that way.

Windows-specific terms

Win+H (Windows voice typing)

Win+H is the keyboard shortcut that opens Windows' built-in voice typing bar on Windows 10 and 11. Free, built into Windows and a sensible first thing to try for everyday dictation. Windows 11 also includes Voice Access, which adds full voice control of the PC along with custom vocabulary and text correction.

Windows Voice Access

Voice Access is Windows 11's hands-free computer control feature — clicking, scrolling, and switching windows by voice, with dictation included. It is built for full computer control by voice. If what you mainly want is comfortable writing, a dedicated dictation tool is usually the better fit.

Windows Speech Recognition (WSR)

Windows Speech Recognition is the older built-in dictation tool that Microsoft has deprecated in favor of Voice Access. If you relied on WSR on an older PC, its retirement is usually the moment to choose a replacement deliberately rather than inherit whichever tool Windows offers next.

Microphone privacy settings

Microphone privacy settings are the Windows switches that decide which programs may use your microphone. If a dictation app hears nothing, this is the first place to look: Settings → Privacy & security → Microphone.

Default microphone

The default microphone is the one Windows hands to programs unless they choose differently. Laptops often have several (built-in, headset, webcam). Picking the right one — close to your mouth, away from fans — improves accuracy more than any setting.

Buying terms

One-time purchase (perpetual license)

A one-time purchase means paying once and owning the software — no monthly or annual fee to keep using it. CaringDictate is $49 once, with updates included. The alternative model, subscriptions, keeps charging for as long as you keep typing.

Subscription

A subscription is paying monthly or yearly for continued access to software — stop paying and it stops working. Reasonable for services with ongoing server costs; harder to justify for a tool running entirely on your own computer. Compare five-year totals before choosing.

Seat (device license)

A seat is one installed copy of a program; a 3-seat license covers three of your computers. CaringDictate's $49 license includes 3 seats — desktop, laptop, and one more.

Free trial

A free trial is a full-featured tryout period before paying — CaringDictate's lasts 7 days with no card required. A trial on your own computer, with your own microphone and your own voice, beats any review or accuracy chart.

Free updates

Free updates means new versions of the software are included in the price you already paid, with no separate upgrade fee. Worth checking before buying any one-time software: some "perpetual" licenses freeze you on the version you bought and charge again for upgrades.

Word cap (usage limit)

A word cap is a limit on how much you may dictate — common on free tiers of subscription dictation services. On-device tools have no reason to meter you; there is no server bill. CaringDictate has no cap.

GPU acceleration

GPU acceleration means using your computer's graphics chip to run the voice model faster than the main processor could. It is why a gaming laptop transcribes noticeably faster. Not required — CaringDictate also runs on ordinary machines with smaller models.

System requirements

System requirements are the minimum computer specs a program needs to run well. For CaringDictate: Windows 10 (64-bit, version 1903) or Windows 11, 4 GB of RAM, and a few GB of free disk space.

The short version

Dictation, voice typing, and speech-to-text all mean the same thing: you talk, the computer types. The choices that actually matter when buying are where your voice is processed (on your computer or on someone's server), how you start and stop it (a comfortable talk button, or a wake word), and how you pay (once, or forever). CaringDictate's answers are: on your computer, any button you like — $49 once, with a 7-day free trial to check it suits your voice first.