Issue link: https://digital.copcomm.com/i/1543070
C I N E M A A U D I O S O C I E T Y. O R G 51 In March 2018, I sat in the Kim Novak Theater at Sony Pictures Studios while two engineers from Sony Corporation of Japan, Toru Nakagawa and Koyuru Okimoto, prepared a binaural demonstration. I had experienced binaural virtualization technologies before and remained skeptical. None had ever convinced me, particularly on the center channel. The center had always been the tell. I sat at the center of the console while they placed tiny microphones just inside my ear canals, which felt a little strange. Sine wave sweeps played through the Novak's speaker system while their software captured the signal from the microphones in my ears. Then I put on a pair of over-the- ear headphones and listened to the same sweeps again. They began playing their demo content in the room. I was about to participate in what would become the universal reaction to this demo. I removed the headphones and the sound disappeared. The sound wasn't in the room. It was in the headphones. Wait. That was in the headphones? I thought there had been a mistake. But there was no mistake. They had completely replicated the sound of the Kim Novak Theater, center channel and all. For the first time, I had been fooled. What I didn't know at the time was that Koyuru had been working on this technology since 2001, and Toru since 2008. What I also didn't know was that the Kim Novak demo had surprised them, too. The system had been developed and tuned in small listening rooms in Tokyo. They had never tested it in a space as large as the Theater and weren't certain it would work. It worked. Three Decades in the Making Sony's work on headphone virtualization stretches back further than most realize. In 1994, the company released the VIP-1000, a two-channel system with head tracking. Over the following two decades, Sony released a succession of digital surround headphone products, including the MDR-DS5000, MDR-DS8000, and others, each expanding channel count and refining the technology. These were consumer products designed to bring theatrical sound into the home. The connection to Sony Pictures Entertainment began formally in 2008, when engineers measured the impulse responses of five dubbing stages on the lot. In 2010, they demonstrated a 96 kHz version of the system. By 2011, the MDR-HW700DS consumer headphones launched with sound modes supervised by SPE, using acoustic measurements captured from actual dubbing stages. The relationship between Tokyo R&D and Culver City SPE Post Production was already well established. But these early systems relied on generic Head-Related Transfer Function (HRTF) profiles. A generic HRTF uses average measurements to approximate how sound reaches human ears, allowing the system to work without measuring each individual listener. For many applications, this is sufficient. For the perfectionists at Sony, it was not. The problem was the center channel. Toru explains the fundamental challenge: Sound from a center speaker arrives at both ears at nearly the same time and level. Unlike sounds from the sides, where timing and level differences between ears provide strong localization cues, the center relies almost entirely on the subtle filtering effects of each person's unique ear shape, head size, and body. A generic profile cannot account for these individual differences. The center image collapsed into the head or smeared across the soundstage. It never sounded like a speaker in front of you. By 2012, the team began pursuing personalized measurement, capturing each listener's individual HRTF rather than relying on averages. This was the turning point. The 2017 and 2018 demonstrations at Sony Pictures, first in the smaller Thalberg Stage, then in the Kim Novak Theater, marked a critical milestone. The technology scaled. What worked in a small Tokyo listening room worked in a large-scale dubbing stage. Following those demos, Yoshi Takashima at Sony Pictures began bridging the gap between the R&D team in Japan and the post-production engineers in Culver City. The sound department had specific requirements: professional quality suitable for critical mixing decisions, Dolby Atmos capability, and the ability to recreate SPE's reference theaters like the Cary Grant. The developers started building software that could meet these demands, incorporating feedback from the mixing and editorial teams. What followed was an intensive period of testing and refinement. Helmed by Tom McCarthy in March 2019, 40 sound designers were scanned and tested on TV production stages. In August, they conducted Dolby Atmos testing. In September, they interviewed TV creators to understand workflow needs. By January 2020, they were testing in the Cary Grant Theater with prototype headphones, measuring profiles for numerous engineers and mixers. A kickoff meeting established a commitment to deploy the technology on three upcoming titles.

