← All articles
musicthe voicesingingacousticsSeptember 15, 20265 min read

How Does the Human Voice Work? Vocal Folds, Resonance and Range

By the BrainSnail editorial team. How these articles are written and checked, and how to tell us when one is wrong.

The sound that comes out of a mouth begins as a buzz, made by two small folds of tissue in the larynx flapping open and shut a hundred or more times a second in the air from the lungs, and it becomes speech or song only in the throat and mouth above, which filter the buzz into vowels, and the tongue, lips and teeth, which chop it into consonants. Every voice is the same instrument, a reed and a resonating tube, and the reasons one person's is recognisable from the next room and another's can fill an opera house are matters of anatomy and of training in equal parts.

The source

The vocal folds are two bands of muscle covered in a loose mucous membrane, stretched across the airway inside the larynx, the box of cartilage behind the Adam's apple. When a person breathes they are held apart; to speak they are brought together, and the air pushed up from the lungs forces them apart, the pressure drops as the air rushes through, and they snap shut again, in a cycle that repeats at the frequency of the note being sung, about 100 times a second for a man's speaking voice, 200 for a woman's, over 1,000 for a soprano's top notes. The result is a buzz, rich in harmonics, and not yet a vowel. Pitch is raised by tightening the folds, using muscles that stretch them as a guitarist stretches a string, and loudness by pushing more air through them, which is why shouting is work for the breath.

The filter

The buzz passes up through the throat, mouth and nose, and that column of air resonates at frequencies set by its shape, reinforcing some harmonics of the buzz and damping others. Move the tongue, jaw and lips and the shape changes, and the pattern of reinforced frequencies, the formants, changes with it; the ear reads the pattern as a vowel. The same buzz shaped by a wide open mouth is ah and by a mouth nearly closed with the tongue high is ee, and the difference between one speaker's ah and another's is the difference in the sizes and shapes of their tubes. The parts and what they do:

  • The lungs and diaphragm: the power supply, whose control is what singers call breath support
  • The vocal folds: the source, setting pitch by their tension and length, and producing a whisper when they are held apart and air rushes through
  • The throat, mouth and nose: the resonator, shaping the buzz into vowels and giving the voice its timbre
  • The tongue, lips, teeth and palate: the articulators, which interrupt the flow to make consonants
  • The false folds, the epiglottis and the ventricles: structures whose adjustment gives the growl, the twang and the operatic ring

Why voices differ

Men's folds are longer and thicker than women's, since testosterone at puberty grows the larynx, which is why the voice breaks and why men's voices are about an octave lower; taller people have longer vocal tracts and lower formants, which is why a large man sounds large. The rest is the fine structure: the exact length and mass of the folds, the shape of the nasal cavity, the habits of a lifetime of speaking in one language and one accent, and the small irregularities in the folds' vibration that give a voice its grain. A voice is as identifiable as a face and for the same reason, that its features are many and their combination is unique, and it changes with age as the folds stiffen and thin, which is why the elderly sound as they do and why a singer's career has a shape.

What singers train

A trained voice is the same instrument used more efficiently. Singers learn to manage the breath so that pressure at the folds stays steady across a phrase; to keep the folds vibrating cleanly through the break between the lower and upper registers, the chest and head voices, where the muscles that set pitch hand over from one to another; and to shape the resonator so that the harmonics of the voice fall where the ear is most sensitive, around 3,000 hertz, which produces the ring that carries an unamplified voice over an orchestra. Range is partly given and partly extended by training, and the classification of voices, from bass through baritone and tenor for men and contralto, mezzo and soprano for women, is by range and weight together. The extremes are wide: a bass may sing down to 65 hertz and a coloratura soprano up past 1,500, and the folds that do both are a couple of centimetres of tissue.

Losing it and keeping it

The folds are delicate, and a voice is lost by using them badly. Shouting, singing without support, smoking and reflux swell and thicken them, and the nodules and polyps of the overworked voice, the singer's nodes that end careers, are calluses on the folds' edges from being slammed together too hard for too long; rest, hydration and retraining cure most of them, and surgery some. The voice also carries the state of the body, which is why a cold, tiredness, fear and grief are audible, and why the stress of a phone call can be heard by the person at the other end; the instrument sits in the throat, between the breath and the brain, and everything that passes through either shows in it.

The takeaway

The human voice is a buzz produced by the vocal folds in the larynx snapping open and shut in the air from the lungs, at a frequency that sets the pitch, filtered into vowels by the resonance of the throat and mouth and cut into consonants by the tongue, lips and teeth. Voices differ by the size and shape of the folds and the tract, and singers train the breath, the passage between registers and the resonance that makes a voice ring, on an instrument that shouting, smoking and strain can damage.

Practise this

Questions from Singing and the Voice

Reading about something is not the same as being able to recall it. These are real questions from the Singing and the Voice unit in our Music track, answers and explanations included. The unit has 118 in total across 23 steps.

  • Sort into groupsLevel 3

    1. Sort each feature into Verse or Chorus.

    Answer: Words change each time = Verse; Tells the story = Verse; Words repeat each time = Chorus; Everyone sings along = Chorus

    Verses change their words and carry the story, while the chorus repeats and everyone sings it.

  • Multiple choiceLevel 2

    2. In call and response singing, one singer sings a line and then...

    • A group answers backcorrect
    • Everyone stops
    • The song ends
    • Nobody listens

    In call and response, a leader sings and a group answers back.

  • True or falseLevel 1

    3. A melody is made of notes sung or played one after another.

    Answer: True

    A melody is a line of notes that follow each other to make a tune.