In this final part, we will interpret mathematically how formants and vocal frequency affect ease of singing, and how changes in the vocal tract alter the formants. This part is independent of Part 2, but if you have not read Part 1, please read it first.
This part uses mathematics and physics at roughly the level of a second- or third-year university course. It involves partial differential equations, perturbation theory, and a concept equivalent to impedance in AC circuits, so please look them up as needed if they are unfamiliar.
1. The Wave Equation
1-1. Derivation
As I mentioned in the introduction to Part 1, I do not understand the physical laws underlying the derivation of the following wave equation very well. It apparently follows from conservation of mass and conservation of momentum, but the LLM's explanation involved topics such as adiabatic processes and went beyond my understanding. As a mathematician, I tend to begin by blindly trusting that an equation is correct, so I am prone to neglecting the physical laws behind it.
As the simplest physical model, treat the vocal tract as a one-dimensional tube with a uniform cross-sectional area. Part 2 also explained that the effect of damping is small enough to be treated as a perturbation. Here too, we assume that energy is conserved without loss and model the vocal tract as a lossless, uniform tube. In the coordinate system from Part 1, the tube extends along the z-axis, but in this part we will use x rather than z for displacement. Define the physical quantities used below as follows:
p(x,t): acoustic pressure, meaning the difference from atmospheric pressure (Pa)
u(x,t): average particle velocity of the air in the vocal tract (m/s)
A: cross-sectional area of the vocal tract (m²)
U(x,t)=Au(x,t): volume velocity of the air in the vocal tract (m³/s)
ρ: density of air (kg/m³)
c: speed of sound in air (m/s)
Considering conservation of mass over an infinitesimal change Δx apparently gives
∂x∂u+ρc21∂t∂p=0.(1)
Conservation of momentum similarly gives the linearized Euler equation
∂x∂p+ρ∂t∂u=0.(2)
Differentiating (1) with respect to t and (2) with respect to x, then eliminating u, gives the partial differential equation for p
∂x2∂2p−c21∂t2∂2p=0.(3)
These equations are the starting point for everything that follows.
1-2. Acoustic Impedance
Equations (1), (2), and (3) from the preceding subsection are all linear in p and u. Therefore, if complex-valued functions p~,u~ satisfying
p(x,t)u(x,t)=Re[p~(x,t)]=Re[u~(x,t)]
solve the wave equation, then p,u are also solutions. This extends the wave equation to complex values. We then define the acoustic impedance Z=Z(x,t) by
Z=U~(x,t)p~(x,t)=Au~(x,t)p~(x,t).
Impedance is complex-valued. Its magnitude ∣Z∣≥0 represents “resistance,” while its argument arg(Z)∈(−π,π] represents the phase difference between pressure and airflow. As discussed in Part 1, the ideal phase relationship has pressure leading the air, particularly with arg(Z)≈2π. We will now consider when this condition is satisfied.
As we will also verify later, resistance is zero in the lossless case. Writing Z=R+iX, R=0 means that the argument of Z depends only on the sign of the reactance X, so arg(Z)=±2π. In reality, however, friction with the throat, air resistance, and other effects dissipate energy, so R>0 and therefore arg(Z)∈(−2π,2π). Note that arg(Z)≈2π is achieved by making the reactance X approach +∞.
2. Formants
2-1. Deriving the Formants
Part 1 described formants as the frequencies that resonate in the vocal tract and are determined solely by its shape. Mathematically, they are frequencies that solve the boundary-value problem for the wave equation below even without an external force.
As in Part 2, assume harmonic oscillations p~(x,t)=P(x)eiωt and u~(x,t)=AU(x)eiωt. This notation is somewhat improper mathematically, but U(x) and U(x,t) are different functions. Let the length of the vocal tract be L, and introduce coordinates with the glottis at the origin x=0 and the lips at the endpoint x=L. The boundary conditions are
no external force and a closed glottis, together with (2), ⟹∂x∂px=0=0; and
release into the atmosphere at the lips ⟹p(L,t)=0.
This gives the boundary-value problem
⎩⎨⎧∂x2∂2p−c21∂t2∂2p=0∂x∂px=0=0,p(L,t)=0.
Substituting p~(x,t)=P(x)eiωt gives the boundary-value problem for the ordinary differential equation in P(x)
{P′′+k2P=0P′(0)=0,P(L)=0,
where k=cω is the wavenumber. The general solution is P(x)=Bcos(kx)+Csin(kx). The boundary conditions imply C=0 and kL=(n−21)π,n=1,2,3,…. Thus, the angular frequencies and solutions are
The frequencies above resonate in the vocal tract without any external force. During actual phonation, however, the lungs supply air and create an external force at x=0, allowing us to produce a voice at any pitch. In this case, we must reconsider the boundary condition ∂x∂px=0=0.
At x=0, the external force acting on the air is proportional to ∂t∂u. If U(0)=Ug=0 in u~(x,t)=AU(x)eiωt, then together with (2), the boundary condition becomes
∂x∂px=0=−ρ∂t∂ux=0=−AiρωUgeiωt.
This gives the boundary-value problem for P
⎩⎨⎧P′′+k2P=0P′(0)=−AiρωUg,P(L)=0.
We now solve it.
The general solution is again P(x)=Bcos(kx)+Csin(kx), and P(L)=0 gives Bcos(kL)+Csin(kL)=0. Suppose the angular frequency ω equals one of the formant angular frequencies ωn. Then cos(kL)=0 implies C=0, which cannot satisfy the boundary condition on P′(0) because Ug=0. Therefore, ω=ωn⟹cos(kL)=0, and
and in particular, the impedance at the glottis x=0 is
Zin=iAρctan(kL).
The impedance diverges when ω=ωn, at a formant angular frequency, and changes sign on either side. The calculation above shows that maximizing the reactance X means keeping the angular frequency ω of the sung pitch slightly below the formant angular frequency ωn.
3. Formant Tuning
3-1. Webster's Horn Equation
Equation (4) makes the formants appear to depend only on the length L of the vocal tract. In actual formant tuning, however, other factors matter, such as raising or lowering the jaw to adjust the size of the mouth. These factors do not appear in (4) because the “uniform” assumption made in 1-1—that the cross-sectional area is constant—is unrealistic. We must therefore treat the cross-sectional area A as a function A(x).
This leads to Webster's horn equation. When the cross-sectional area is not uniform, equation (1), derived from conservation of mass, becomes
A1∂x∂(Au)+ρc21∂t∂p=0.(5)
Equation (2), derived from conservation of momentum, remains valid. Differentiating (5) with respect to t—noting that A does not depend on t—and then using (2) to eliminate u gives Webster's horn equation
A1∂x∂(A∂x∂p)−c21∂t2∂2p=0.(6)
Assuming a harmonic solution p(x,t)=P(x)eiωt gives the ordinary differential equation for P
dxd(A(x)dxdP)+c2ω2A(x)P=0.(7)
We will use this equation with the boundary conditions
P′(0)=0,P(L)=0
to determine how changes in A(x) affect the formants.
3-2. Perturbation Theory
Webster's horn equation cannot be solved explicitly for a general A(x). We therefore treat the uniform equation (3), with A(x)=A0, as the unperturbed case. A nonuniform tube A(x) is viewed as the uniform tube A0 with a perturbation δA(x), and we use (7) to determine how that perturbation affects the formants.
Let 0<ε≪1 control the magnitude of the perturbation, and write A(x) as
A(x)=A0+εδA(x).
Similarly, expand ωn and Pn(x) as perturbations of their unperturbed solutions:
At last, the part where we thoroughly abuse the equations is over. All that remains is to consider how changing the vocal tract changes the formants. The only physical quantities we can consciously alter are essentially the length L of the vocal tract—by raising or lowering the larynx, protruding the lips, and so on—and the perturbation δA(x) to its cross section.
The effect of L is straightforward. Recalling (4),
Fn=4L(2n−1)c,n=1,2,3,…,
every formant Fn is inversely proportional to L and therefore decreases monotonically with it. In practice, singing with pursed lips is not realistic outside of training exercises, so it appears that L is adjusted by raising or lowering the larynx.
The effect of a cross-sectional change δA(x) is not so simple. Recalling (8),
The factor cos(L(2n−1)πx) in the integrand changes sign over [0,L], and its behavior also depends on n=1,2,3,…. We therefore cannot draw a simple conclusion that increasing A(x), meaning δA(x)>0, always raises or lowers a formant. One clear result, however, is that lowering the jaw and opening the mouth wide means δA(x)>0 near x≈L. In this region, cos(L(2n−1)πx)≈−1 regardless of n, so the formants rise. Conversely, closing the mouth somewhat lowers the formants.
This explains why inexperienced vocalists like me tend to open their mouths wide and shout when singing high notes. Vowel modification, the quintessential formant-tuning technique, instead teaches us to pronounce a wide-open “ah” slightly closer to “eh” or “oo,” with the mouth somewhat more closed. This lowers the formants. The LLM claims that “true” mixed voice is achieved by matching the second formant F2, rather than the first formant F1, to the fundamental frequency f0. Some LLMs explained that harmonics nf0 should be matched to F2 or F3, but I honestly do not know what is correct.
Conclusion
Thank you to everyone who read this far. You have now reached the same point I occupied while writing: “I sort of understand the principles of mixed voice, but I still cannot actually sing anything.” The next step is to put the theory into practice and turn skillful maneuvers of the vocal organs into muscle memory. The necessary vocal training will presumably require steady daily practice. If this series can do even a little to keep someone from giving up before seeing results and wondering, “Why do I have to do this?”, I will be delighted.
Loading comments.