SPEECH CHIP COLLECTION

MM54104 (Digitalker)

National Semiconductor’s Digitalker speech processor reconstructed compressed human speech from external vocabulary ROMs.

DT1000 / MA6100 Digitalker Evaluation Board — photograph 3.

Overview

National Semiconductor’s Digitalker speech processor reconstructed compressed human speech from external vocabulary ROMs.

Image gallery

A real speaker, compressed into silicon

The MM54104 is the speech processor at the heart of National Semiconductor’s Digitalker system. Its vocabulary resides in external memory, not inside the processor. The familiar DT1050 kit combines one MM54104 with two MM52164 speech ROMs, SSR1 and SSR2. A filter, amplifier, and loudspeaker complete the system. National’s December 1980 documentation describes a standard male voice while explaining that other encoded vocabularies can reproduce female and children’s voices. The voice belongs to the supplied speech data; the processor does not inherently have only one voice.

The Mozer approach

Digitalker belongs to the speech-compression tradition associated with Forrest S. Mozer, also represented by the earlier S14001A. Mozer and Richard P. Stauduhar’s patent describes reducing recorded speech through waveform manipulation, selective repetition, silence coding, and delta modulation. The aim is to preserve recognizable speech while storing much less information. This differs from the SC-01’s analog vocal-tract model and the T6721A’s PARCOR filter-based reconstruction. Digitalker reconstructs encoded waveform material derived from a speaker rather than accepting English text or an unrestricted phoneme command stream.

How the compressed speech is reconstructed

MAME’s documented reverse engineering gives a useful view of the decoder. For voiced sounds, its model rebuilds waveform periods from compact adaptive differential codes, using symmetry, reversal, and—in some modes—zero-filled portions. Pitch and amplitude information control how those periods are played and repeated. Unvoiced sounds use another decoding mode, and silent intervals are represented by duration information. The ROM contains tables pointing to sequences of sound segments, so one command can retrieve an entire word or phrase. These details come from emulator research; its source explicitly identifies an uncharacterized mode, so it should not be treated as a complete manufacturer specification.

What the standard vocabulary actually contains

The DT1050 datasheet lists 144 addressable expressions: 137 words, two tones, and five silence durations. Its address table runs from 0 to 143; the highest address is not the vocabulary count. The spoken material includes letters, numbers, measurement terms, directions, and warning words. The numbered pauses help join individual words into more convincing sentences. Vocabulary counts in later literature can differ because some totals count words while others include tones, pauses, or introductory phrases.

More than one vocabulary kit

National offered several ROM sets for different uses. DT1051 demonstrated complete phrases and different speakers, including a child’s voice and a bassoon example. DT1052 supplied digits and “point” for instruments and numerical readouts. DT1056 supplied a second general vocabulary; DT1057 sold its SSR5/SSR6 ROM pair separately. National’s August 1981 sheet lists 131 entries for that set and describes expansion of DT1050 using additional selection circuitry. Whole recorded phrases preserve their original timing better than sentences assembled one word at a time.

Creating custom speech ROMs

Custom vocabulary did not necessarily mean commissioning an entirely new recording. National’s DTSW-500 Vocabulary Selection System supplied CP/M software and an archive of 500 encoded male-voice words on two eight-inch floppy disks. A designer wrote a vocabulary list, checked that its words existed in the archive, compiled a work file, and built a binary ROM image for programming PROMs. This selected and assembled already encoded speech; it was not a general converter for typing any English word or importing arbitrary audio. Creating genuinely new speech material required the separate recording and encoding process. National described the archive format as expandable to additional standard or custom words.

Talking to the processor

The host presents an eight-bit word or phrase number on SW1–SW8 and selects the chip with CS low. With CMS low, a write-strobe cycle resets the interrupt and starts speech on the rising edge of WR. INTR goes high when the sequence finishes, allowing the host to send the next word. CMS high instead clears the interrupt without starting another sequence. A new speech command can interrupt the current one immediately. The processor supplies a fourteen-bit ROM address bus and reads speech data through a separate eight-bit bus; its directly addressable speech memory is 128 kilobits, or 16 kilobytes.

Timing, pauses, and word-building tricks

The host can chain vocabulary entries to announce measurements, times, or status messages. National recommends deliberate pauses to improve phrase rhythm rather than simply firing words back-to-back. Experimenters can also interrupt one word with another to splice fragments, or gate the audio to retain only part of a word. The result depends on careful timing and the available recordings; this does not turn Digitalker into an unrestricted text-to-speech system. The chip also supports a manually operated switch interface with debounce circuitry, so a processor is not essential for a simple talking demonstrator.

The supporting electronics

The MM54104 uses a 7–11 V supply and a nominal 4 MHz clock. Its control interface is TTL-compatible, while the vocabulary ROM supply must follow the ROM’s own specifications. The analog speech output needs external filtering and amplification. ROMEN can control external switching circuitry to reduce ROM power consumption. These characteristics suited fixed-message products such as instruments, clocks, alarms, telephone equipment, and talking aids. Modern hobby projects keep the original processor but substitute programmed parallel flash memory for banks of vintage ROMs.

Technical reference

FeatureDetails
ManufacturerNational Semiconductor
ProcessorMM54104 / MM54104N; 40-pin DIP
Speech methodMozer encoded speech; specialized waveform reconstruction
VocabularyExternal speech ROMs or compatible programmed memory
Direct speech memory128 kbit / 16 KiB
Host selection8-bit word/phrase command
ClockNominal 4 MHz
Processor supply7–11 V
AudioAnalog output; external filter and amplifier

Digitalker — numbers and alphabet

Hear the Digitalker speak numbers and letters.

Documentation and downloads

Machines using this chip

Further references