Votrax · Analog formant / phoneme synthesis · 1980

SC-01 / SC-01A

The phoneme voice behind talking robots, home-computer accessories, and some of the most memorable machines of the early 1980s.

SC-01 / SC-01-A speech chip

Overview

The Votrax SC-01 creates speech by combining small speech sounds rather than replaying recordings of complete words. Introduced at the beginning of the 1980s, this 22-pin CMOS chip brought an electronic model of the vocal tract into a single integrated circuit. With suitable programming, its small repertoire could form new words, sentences, and even deliberately nonsensical voices. The SC-01 and revised SC-01A became recognizable voices in computer peripherals, robotics, arcade games, and pinball.

Compare the Votrax family, interfaces, and voices →

Image gallery

SC-01 versus SC-01A: a refined voice

The SC-01A (also marked SC-01-A) is a voice refinement of the SC-01, retaining its 22-pin package, basic interface, and phoneme-code repertoire. After examining both chips’ internal ROMs, reverse-engineer Jonathan Gevaryahu reported changes to parameters for a few phonemes, apparently intended to improve sound quality and remove some DC bias. MAME likewise uses the same synthesizer implementation with separate phoneme ROM images. The documented distinction is therefore in the sound parameters, rather than a new programming architecture. Their voices need not sound identical.

From a screw manufacturer to a talking chip

Votrax grew out of the speech work of Richard T. Gagnon and the vocal division of Federal Screw Works. Its earlier synthesizers used larger assemblies; the SC-01 condensed phoneme synthesis into a single chip. The 1980 datasheet still identifies Votrax as a division of Federal Screw Works in Troy, Michigan. This unusual industrial background helped produce a voice that would become familiar far beyond the factory.

Unlimited vocabulary does not mean text-to-speech

The chip has 64 selectable codes, including the PA0 and PA1 pauses and a STOP code. They are selected through six data inputs, P0–P5. Several codes represent related versions of a speech sound, so the repertoire should not be mistaken for 64 different English letters. Words are assembled from pronunciation, not spelling: a program must choose the appropriate sound sequence and pauses. The SC-01 has no built-in English dictionary or spelling parser. A system such as the Type N Talk supplies the text-conversion software outside the chip. A fixed-message product can instead store prepared phoneme sequences in an ordinary external ROM.

An electronic vocal tract

The SC-01 is an analog formant synthesizer. Voiced excitation supplies the periodic energy associated with vocal-fold vibration, while noise excitation supplies the hiss needed for unvoiced consonants. Controlled filters shape those sources into speech sounds. A formant is a resonance that helps distinguish one vowel from another; moving the filter settings changes the effective shape of this electronic vocal tract. The internal phoneme memory holds control parameters rather than recordings of someone speaking. Modern reverse engineering, reflected in MAME’s circuit simulation, models separate voice and noise paths, multiple filter stages, closure control, and a final output filter. Its 512-byte phoneme ROM is a compact collection of sound recipes, not a vocabulary of stored words.

Why the transitions matter

In speech, the mouth keeps moving as one sound becomes another. Simply placing isolated sounds next to one another can lose those movements. The Votrax design changes its control values progressively between phonemes, producing intermediate states as the electronic vocal tract moves toward its next sound. Gagnon’s patent describes digital transition circuitry that gradually updates stored parameter values. This is an important part of the chip’s character and intelligibility, although it does not make its voice a natural human recording. Pronunciation, timing, and stress still depend on the phoneme sequence supplied by the programmer.

Four pitch levels—and a clock that changes the whole voice

Two inflection inputs, I1 and I2, select four pitch levels for voiced sounds. Software can change them through an utterance to introduce emphasis or a rising or falling pitch. The master clock offers another way to alter the voice. The datasheet specifies 720 kHz for standard phoneme timing and provides both an internal RC oscillator and an external-clock option. Reducing the clock frequency lowers the audio frequencies and lengthens phonemes; increasing it raises the voice and speeds it up. Clock adjustment therefore changes speech rate and voice character together. It is distinct from selecting the four inflection levels.

Programming the SC-01

The controller places a six-bit phoneme code on P0–P5 and pulses STB; the data is latched on the rising edge. The acknowledge/request output, A/R, goes low to acknowledge the request and returns high when the phoneme’s timing interval has finished. A controller can poll it or use it to request the next code through an interrupt. Six phoneme lines, a strobe, and the request line provide the basic eight-signal interface when inflection is held fixed. Two more controlled signals allow programmed pitch changes. The chip is not a sentence buffer: the surrounding electronics or software must keep feeding it codes. Its datasheet advertises only about 70 bits per second for continuous speech, a strikingly small control-data requirement.

From stored messages to a conversational peripheral

The manufacturer illustrates several ways of building a talking product. In the simplest, a counter steps through phoneme codes held in ROM, paced by A/R. Multiple messages can occupy separate ROM blocks. A microprocessor provides a more flexible arrangement, choosing phonemes as a program runs and converting text when suitable software is available. An external message ROM is therefore optional storage for what to say. It is different from the chip’s internal phoneme-parameter ROM, which defines how its sounds are synthesized.

The output is analog

AO is an analog audio output with a DC bias. It needs an appropriate coupling and amplifier circuit to drive a loudspeaker. The chip also supplies AF, an audio-feedback connection, and CB, a current source used in the datasheet’s transistor amplifier arrangements. This is a mixed-signal device: digital inputs command an analog speech-generating circuit. Its audio output should not be described as a digital PWM speaker signal. Likewise, latched inputs described as 5-volt logic compatible do not mean the complete chip is a modern 5-volt-only part; the supply, inflection inputs, and A/R output must be handled according to the relevant electrical specifications.

The later SC-02 / SSI-263

The SC-02, also known as the SSI 263, belongs to a later generation. It adds programmable registers and much more control over speech parameters, with a different package and pinout. It is a related voice technology rather than a drop-in replacement for SC-01 or SC-01A. The SP0256-AL2 and TMS5220 use different synthesis architectures; their encoded speech data and programming interfaces are not interchangeable with Votrax phoneme codes.

A voice with a personality

The same chip family could speak a carefully programmed sentence or become a character voice. Type N Talk translated computer text into phoneme sequences. The HERO speech card gave a robot an electronic voice. Q*bert used speech hardware for expressive gibberish, while Black Hole used intelligible spoken lines. The samples below let you hear these very different applications. The short “Defender” and “Robots” clips are spoken-word demonstrations; their titles alone do not identify a particular game or robot.

Building and preserving the voice

Keeping an SC-01 system speaking requires its complete signal chain: valid phoneme and strobe timing, the correct clock, an appropriate supply, and the audio amplifier. The phonetic dictionary is especially useful when reproducing a historical phrase because it supplies prepared pronunciation sequences. Steve Ciarcia’s September 1981 BYTE project and the two-part Microvox project in September and October 1982 show how enthusiasts turned the chip into practical computer speech hardware. His March 1984 third-generation project belongs to the later SSI 263. Chip Talk and the Circuit Cellar books provide further routes into hands-on speech synthesis.

The SC-01 interface at a glance

ConnectionPurpose
P0–P5Six-bit phoneme selection
STBRising-edge data latch
A/RAcknowledge / request for the next phoneme
I1, I2Four voiced-pitch levels
MCRC, MCXInternal RC oscillator or external clock
AOAnalog audio output
AF, CBConnections for the recommended amplifier circuits

Hear the Type N Talk

“The juice of lemons makes fine punch. A box was thrown beside the parked truck.”

Hear Q*bert

Q*bert’s expressive, speech-like gibberish.

Hear Ugg

Ugg’s complaining voice from Q*bert.

“Defender”

A short spoken-word demonstration.

“Robots”

A short spoken-word demonstration.

Hear Black Hole

“Do you dare to enter the black hole?” — Gottlieb, 1981.

Documentation and downloads

Machines using the SC-01 family

These exhibits use SC-01 or SC-01A hardware. Some product families changed synthesizers between models; their individual pages explain those differences.

Continue exploring

Further references