SPEECH CHIP COLLECTION
MM54104 (Digitalker)
National Semiconductor’s Digitalker speech processor reconstructed compressed human speech from external vocabulary ROMs.

Overview
National Semiconductor’s Digitalker speech processor reconstructed compressed human speech from external vocabulary ROMs.
Image gallery


Find a Digitalker
Search for Digitalker processors, vocabulary ROMs, and speech kits.
Search eBay for DigitalkerA real speaker, compressed into silicon
The MM54104 is the speech processor at the heart of National Semiconductor’s Digitalker system. Its vocabulary resides in external memory, not inside the processor. The familiar DT1050 kit combines one MM54104 with two MM52164 speech ROMs, SSR1 and SSR2. A filter, amplifier, and loudspeaker complete the system. National’s December 1980 documentation describes a standard male voice while explaining that other encoded vocabularies can reproduce female and children’s voices. The voice belongs to the supplied speech data; the processor does not inherently have only one voice.
The Mozer approach
Digitalker belongs to the speech-compression tradition associated with Forrest S. Mozer, also represented by the earlier S14001A. Mozer and Richard P. Stauduhar’s patent describes reducing recorded speech through waveform manipulation, selective repetition, silence coding, and delta modulation. The aim is to preserve recognizable speech while storing much less information. This differs from the SC-01’s analog vocal-tract model and the T6721A’s PARCOR filter-based reconstruction. Digitalker reconstructs encoded waveform material derived from a speaker rather than accepting English text or an unrestricted phoneme command stream.
How the compressed speech is reconstructed
MAME’s documented reverse engineering gives a useful view of the decoder. For voiced sounds, its model rebuilds waveform periods from compact adaptive differential codes, using symmetry, reversal, and—in some modes—zero-filled portions. Pitch and amplitude information control how those periods are played and repeated. Unvoiced sounds use another decoding mode, and silent intervals are represented by duration information. The ROM contains tables pointing to sequences of sound segments, so one command can retrieve an entire word or phrase. These details come from emulator research; its source explicitly identifies an uncharacterized mode, so it should not be treated as a complete manufacturer specification.
What the standard vocabulary actually contains
The DT1050 datasheet lists 144 addressable expressions: 137 words, two tones, and five silence durations. Its address table runs from 0 to 143; the highest address is not the vocabulary count. The spoken material includes letters, numbers, measurement terms, directions, and warning words. The numbered pauses help join individual words into more convincing sentences. Vocabulary counts in later literature can differ because some totals count words while others include tones, pauses, or introductory phrases.
More than one vocabulary kit
National offered several ROM sets for different uses. DT1051 demonstrated complete phrases and different speakers, including a child’s voice and a bassoon example. DT1052 supplied digits and “point” for instruments and numerical readouts. DT1056 supplied a second general vocabulary; DT1057 sold its SSR5/SSR6 ROM pair separately. National’s August 1981 sheet lists 131 entries for that set and describes expansion of DT1050 using additional selection circuitry. Whole recorded phrases preserve their original timing better than sentences assembled one word at a time.
Creating custom speech ROMs
Custom vocabulary did not necessarily mean commissioning an entirely new recording. National’s DTSW-500 Vocabulary Selection System supplied CP/M software and an archive of 500 encoded male-voice words on two eight-inch floppy disks. A designer wrote a vocabulary list, checked that its words existed in the archive, compiled a work file, and built a binary ROM image for programming PROMs. This selected and assembled already encoded speech; it was not a general converter for typing any English word or importing arbitrary audio. Creating genuinely new speech material required the separate recording and encoding process. National described the archive format as expandable to additional standard or custom words.
Talking to the processor
The host presents an eight-bit word or phrase number on SW1–SW8 and selects the chip with CS low. With CMS low, a write-strobe cycle resets the interrupt and starts speech on the rising edge of WR. INTR goes high when the sequence finishes, allowing the host to send the next word. CMS high instead clears the interrupt without starting another sequence. A new speech command can interrupt the current one immediately. The processor supplies a fourteen-bit ROM address bus and reads speech data through a separate eight-bit bus; its directly addressable speech memory is 128 kilobits, or 16 kilobytes.
Timing, pauses, and word-building tricks
The host can chain vocabulary entries to announce measurements, times, or status messages. National recommends deliberate pauses to improve phrase rhythm rather than simply firing words back-to-back. Experimenters can also interrupt one word with another to splice fragments, or gate the audio to retain only part of a word. The result depends on careful timing and the available recordings; this does not turn Digitalker into an unrestricted text-to-speech system. The chip also supports a manually operated switch interface with debounce circuitry, so a processor is not essential for a simple talking demonstrator.
The supporting electronics
The MM54104 uses a 7–11 V supply and a nominal 4 MHz clock. Its control interface is TTL-compatible, while the vocabulary ROM supply must follow the ROM’s own specifications. The analog speech output needs external filtering and amplification. ROMEN can control external switching circuitry to reduce ROM power consumption. These characteristics suited fixed-message products such as instruments, clocks, alarms, telephone equipment, and talking aids. Modern hobby projects keep the original processor but substitute programmed parallel flash memory for banks of vintage ROMs.
Technical reference
| Feature | Details |
|---|---|
| Manufacturer | National Semiconductor |
| Processor | MM54104 / MM54104N; 40-pin DIP |
| Speech method | Mozer encoded speech; specialized waveform reconstruction |
| Vocabulary | External speech ROMs or compatible programmed memory |
| Direct speech memory | 128 kbit / 16 KiB |
| Host selection | 8-bit word/phrase command |
| Clock | Nominal 4 MHz |
| Processor supply | 7–11 V |
| Audio | Analog output; external filter and amplifier |
Digitalker — numbers and alphabet
Hear the Digitalker speak numbers and letters.

