- Preprint Article
- 10.2139/ssrn.6437130
Predicting Cognitive Performance and Brain Structure From Hearing Assessments in Healthy Aging Males: A Longitudinal Cohort Study
- Jan 01, 2026
- SSRN Electronic Journal
- Kamine Julie Jacobsen + 11 more +11
Publications from 2021 to 2026
Showing 4 of 4 papers
Predicting Cognitive Performance and Brain Structure From Hearing Assessments in Healthy Aging Males: A Longitudinal Cohort Study
Adult Users of the Oticon Medical Neuro Cochlear Implant System Benefit from Beamforming in the High Frequencies
The Oticon Medical Neuro cochlear implant system includes the modes Opti Omni and Speech Omni, the latter providing beamforming (i.e., directional selectivity) in the high frequencies. Two studies compared sentence identification scores of adult cochlear implant users with Opti Omni and Speech Omni. In Study 1, a double-blind longitudinal crossover study, 12 new users trialed Opti Omni or Speech Omni (random allocation) for three months, and their sentence identification in quiet and noise (+10 dB signal-to-noise ratio) with the trialed mode were measured. The same procedure was repeated for the second mode. In Study 2, a single-blind study, 11 experienced users performed a speech identification task in quiet and at relative signal-to-noise ratios ranging from −3 to +18 dB with Opti Omni and Speech Omni. The Study 1 scores in quiet and in noise were significantly better with Speech Omni than with Opti Omni. Study 2 scores were significantly better with Speech Omni than with Opti Omni at +6 and +9 dB signal-to-noise ratios. Beamforming in the high frequencies, as implemented in Speech Omni, leads to improved speech identification in medium levels of background noise, where cochlear implant users spend most of their day.
Read moreA binaural short time objective intelligibility measure for noisy and enhanced speech
Objective intelligibility measures are increasingly being used to assess the performance of speech processing algorithms, e.g. for hearing aids. It has been shown that the short time objective intelligibility (STOI) measure yields good results in this respect. In this paper we propose a binaural extension of the STOI measure, which predicts binaural advantage using a modified equalization cancellation (EC) stage. The proposed method is evaluated for a range of acoustic conditions. Firstly, the method is able to predict the advantage of spatial separation between a speech target and a speech shaped noise (SSN) interferer. Secondly, the method yields results comparable to the monaural STOI measure when presented with noisy speech processed by ideal time-frequency segregation (ITFS). Finally, the method also performs well when presented with a selection of different acoustic conditions combined with beamforming as used in hearing aids.
Read moreEnhancing music with virtual sound sources
For many people, listening to music is an important part of life. Most often the music is recorded and played on a CD player, the radio, the television, an mp3 player, or a computer. Listening to music from such devices was long out of reach for hearing aid users. But recently, the development of devices, such as the Oticon Streamer, that can send music wirelessly to hearing aids enables people to enjoy listening to music directly in their hearing aids with a good signal-to-noise ratio. However, listening to music sent directly to hearing aids is not optimal. Specifically, the sound image appears to be inside the listener's head. This is referred to as “in-the-head locatedness.”1 When the signal is the same at both ears (monophonic), the listener perceives it as being in the middle of his or her head. When the signal is stereophonic, the sound is perceived as being on a line between the ears. By changing the level of the signal in either ear, the sound can be moved between the ears. This is referred to as “lateralization of the sound image.” Thus, with a stereophonic signal the sound image can be lateralized, but it is still perceived as being inside the head. Users generally experience this as unpleasant and unnatural since it is not what occurs in real-life listening, where sound sources are placed at a distance in the space around the listener. Therefore, it is desirable to enhance this sound image to make it more natural and pleasant to listen to. The problem of in-the-head locatedness also occurs when people listen to stereo music through ordinary headphones. This is because stereo music is designed to be played through two loudspeakers placed in front of a listener in a room. Specifically, the loudspeakers have to be placed at ±30° if the listener is to perceive the correct stereo image. With this setup, the sound image is perceived as between the loudspeakers, which creates the correct spatial sound stage. Since the sound image is out in the room, not in the head, normal binaural hearing can be used to localize the sound. Since recorded music is intended to be listened to through loudspeakers in a room, it is desirable to simulate this when listening through hearing aids (or headphones). This can be done by filtering the hearing aid signals similarly to how they are normally filtered by loudspeakers in a room. This method, known as “binaural synthesis,”2 has been used in developing a new algorithm called Music Widening, which is implemented in Oticon's latest premium wireless hearing aids, Agil. This article explains the principles behind the algorithm and its implementation and reports the results of a clinical listening test. CREATING A VIRTUAL ENVIRONMENT The central idea of binaural synthesis is that the sound pressure at the listener's eardrums contains all acoustical information about a sound event, including spatial aspects. Therefore, if the sound pressure at the eardrums is the same on two occasions, the listener should experience the same auditory event. This idea is not new. In 1863, Helmholtz wrote, “If the motions of the particles of air in the aural passage are the same on two different occasions, the ear will receive the same sensation, whatever the origin of the motions.”3 Therefore, by carefully controlling the signal in the ear canal, it is possible to simulate an auditory event that does not correspond to a real sound event. Thus, binaural synthesis can be used to create sound sources in a virtual environment. To understand how, consider the sound at the ears of a listener coming from a source in a non-reflecting (anechoic) room. When sound is played through the source, the incoming sound wave interacts with the torso, head, and pinnae of the listener before entering the ear canals. This interaction of the sound with the body can be described by the head-related transfer functions (HRTFs), which are defined as the sound pressure in the ear as a function of direction. HRTFs depend heavily on the angle of incidence of the sound wave with respect to the listener. As an example, see the pair of HRTFs in Figure 1. They have been measured in an anechoic room with Epoq XW hearing aids on an artificial head. A loudspeaker was placed directly to the left of the head, i.e., in the horizontal plane at 90°. The HRTFs are shown as impulse responses, as a function of time, separately for the left and right ears. The figure shows that the sound arrives at the left ear before it arrives at the right ear. Furthermore, it is seen that the amplitude (sound level) is relatively high at the left ear, but low at the right ear due to the head shadow effect. Similar left/right pairs of HRTFs can be shown for every direction on the sphere around the head.Figure 1: HRTFs for the left (blue) and right (green) ears for a sound source on the side of the head.HRTFs can be used to create virtual sound sources. This is done by filtering a signal with a left/right pair of HRTFs from a certain direction. When the resulting left and right ear signals are played through hearing aids, the listener perceives the sound as coming from a point in space that corresponds to the chosen direction. By selecting another pair of HRTFs one can change the position of the virtual sound source in any direction or even create a moving sound source. Furthermore, it is possible to create more than one virtual sound source at a time. This is done for each additional virtual sound source simply by filtering a signal with a pair of HRTFs and adding the result to the signal played in the hearing aids. In the simulation described above there are no room reflections. But to simulate a loudspeaker in a reflective environment, the characteristics of the room and positions of the sound source and listener in the room must be available.4 From this information one can derive a “room impulse response.” In general, a room impulse response can be divided into three components with slightly different characteristics: the direct sound, early reflections, and late reflections. It is the direct sound that arrives first at the listener's ears without interaction with the room. The early reflections are contributions of sound that arrives at the listener after it has bounced off the walls of the room. The late reflections reach the listener even later and have a more unpredictable character because they consist of many sound contributions from all directions. Figure 2 shows a schematic of the components of a room impulse response. Notice that the room reflections are lower than the direct sound and that their amplitudes gradually decrease with time. This is because sound is absorbed by the walls of the room.Figure 2: The different components of a room impulse response.In binaural synthesis the room impulse response is combined with HRTFs to simulate a virtual sound source in a room. This is done by replacing every sound contribution (the direct sound as well as every individual reflection) by a pair of HRTFs, since each contribution comes from a different angle. The particular pair of HRTFs depends on the direction from which the sound arrives at the listener. The directions and times of arrival of the reflections can be calculated by means of a mirror-image model of the room. So, every reflection is slightly different at the two ears, in that the times and amplitudes are as found in the particular HRTFs for the corresponding direction. Furthermore the reflections are gradually more damped as they hit more walls. Implementing this binaural room impulse response ensures that the acoustical cues at the two ears are correct. The room reflections make a positive contribution to the perception of the sound source. First, they help determine the distance to the sound source. For this, the time between the direct sound and the early reflections, as well as the direct-to-reverberant ratio (the level of the direct sound with respect to that of the reflections), is very important. Manipulating the reflections can move the virtual sound source closer or farther from the listener. The perception of the direction of the sound is determined by the direct sound. This phenomenon is known as the “precedence effect.”1 Furthermore, it is well known from concert hall acoustics that reflections help create the feeling of spaciousness.6 Specifically, early lateral reflections (from the sides) are very important. Thus, by using such discrete reflections, it is possible to create sound images that appear to be broad. Notice that the direct sound and the reflections can come from all directions around the listener. Thus, HRTFs from all directions on the sphere around the head are in principle needed for the simulation. By using such HRTFs, one can place the virtual sound sources naturally in the acoustical space around the listener. The sound does not appear to be inside the head even though hearing aids are used. This is in contrast to what happens when a room effect is created with traditional stereo reverberation processing.7 Applying reverberation to a “dry” voice or musical instrument signal gives it a more pleasant sound, but does not create out-of-head localization. To summarize, binaural synthesis can be used to create an auditory virtual environment. The advantage of this over other methods of sound reproduction is that the listener has the experience of being in the virtual environment. This allows the listener to use the full potential of his auditory system in everyday life. THE MUSIC WIDENING ALGORITHM The goal of the Music Widening algorithm is to give the listener a spacious and pleasant sound image when listening to music through the Oticon Streamer. The idea is to create a virtual listening room with virtual sound sources based on the principles discussed above. Since the algorithm was designed for hearing aids, it was important that it be as efficient as possible to minimize the use of processing power. The principles used in the Music Widening algorithm can be described by considering a sound source (loudspeaker) in front of a listener in a simplified room with only two side walls. The sound travels directly from the source to the listener, and is also reflected off the walls. These reflections can be modeled as sound sources that are flipped across (or mirrored by) the walls. These mirrored sound sources can be considered copies of the original source and they produce sound at the same time. The idea is shown in Figure 3. Notice that the sound source is slightly to the left of the listener and that the distances to the two walls differ. This causes the sound from the two reflections to arrive at the listener at slightly different times.Figure 3: Sound reflections are modeled as mirror images of the sound source.This listening situation can be simulated by binaural synthesis in which virtual sound sources replace each source (the direct sound as well as the reflections). Each virtual sound source is modeled individually with a pair of HRTFs from the corresponding direction. This ensures that the differences between the two ears are correct for every sound source. The concept is shown in Figure 4.Figure 4: Each virtual sound source is implemented with a pair of HRTFs.Since the sources are different distances from the listener they are attenuated differently and arrive at different times. The further away a sound source, the more the sound is attenuated and the later it arrives. In addition, the reflections are damped due to the partial absorption of sound by the walls. By taking all this into account the contributions of sound at the ears can be calculated as impulse responses. The left and right impulse responses due to the three virtual sound sources are shown in Figure 5.Figure 5: Left and right impulse responses due to the virtual sound sources.The example described is very simplified since it has only two reflections. However, for subsequent reflections in a room the principles remain the same. As sound is reflected from all walls (including the floor and the ceiling) multiple times, the sound arrives from all directions. By simply repeating the procedure of creating mirror images and replacing them with virtual sound sources, one can create a complete binaural room impulse response. The more reflections that are created, the more precise and realistic the simulation will be. In a hearing aid, however, it is not possible to simulate all reflections, and the implementation has to be limited to a certain number of virtual sound sources. Another simplification in the example is that there is only one sound source in the room, i.e., only one channel (mono) is available. The simulation can, however, easily be expanded if another channel is available, as for stereo. The same processing as described above is simply repeated for the second channel. If only mono is available the sound source can be placed directly in front of the listener, whereas when stereo is available two sources can be created at ±30° with respect to the listener. The Music Widening algorithm has many adjustable parameters, including the direction and distance to the sound source (direct sound). The room size and the distances to the walls can also be adjusted. The absorption properties of the walls can be changed and one can set the number of reflections to implement in the hearing aids. All these parameters were fine-tuned to find the optimal balance between spatial perception and processing resources. Since the preference potentially depended on the signal (speech or music), many signals were used for the fine-tuning. By optimizing all room parameters, the final left and right impulse responses that are implemented in the hearing aids were selected and the parameters were fixed. However, the dispenser can adjust the overall level of the reflections in accordance with a patient's individual preferences. A LISTENING EXPERIMENT The Music Widening algorithm was tested in a listening experiment. The goal was to determine what level of reflections (referred to as the Level parameter) listeners prefer for different kinds of signals. The 11 participants had hearing losses ranging from mild to severe. Their hearing aids vent sizes varied from 1 mm to completely open. All listeners were binaurally fitted with Oticon Agil hearing aids and each hearing loss was compensated individually. The hearing aids were connected to the Oticon Streamer so sound could be played in the hearing aids. Four signals were used as test stimuli: Speech (male voice), Mixed (male voice and background music), Music 1 (jazz/pop music), and Music 2 (classical orchestra). The participants' task was to listen to two signals alternately and indicate which they preferred. The test was performed by AB-comparisons of stimuli. On each trial, the listeners were presented with one of the four signals. They could freely switch between A and B by pressing on a touch-screen. By doing so they set the Level parameter in the hearing aids either to OFF (no processing), LOW, MED, or HIGH. Thus, six comparisons were made for each signal. The paired-comparison data obtained from the listening experiment were analyzed by the Bradley-Terry-Luce (BTL) model8,9 to obtain preference probabilities. The BTL-scale is normalized so the score can be interpreted as the probability that a given condition is preferred over the other conditions. Figure 6 shows the results for all signals and listeners. The error bars indicate 95% confidence intervals, i.e., conditions with non-overlapping error bars are significantly different.Figure 6: Preferred level of Music Widening for different signals (11 hearing-impaired listeners).The results show that Music Widening with its maximum setting (HIGH) is preferred independent of the signal. This outcome is favorable in terms of implementation, since it means the level need not be adjusted individually depending on the signal. However, the algorithm was developed mainly to improve music listening, hence the name Music Widening. After listeners completed the experiment they were asked to describe the differences they heard, as well as explain what they based their preference on. The interview included questions about sound quality, comfort, speech intelligibility, loudness, spectral balance, and overall preference. Listeners said that they generally chose the sound that they experienced as most spacious. This term generally referred to sound that is the most open, gives the broadest sound image, is most airy, feels like sound in a big room, gives the most reverberation, and gives the best spatial perception. Generally they disliked sounds that they experienced as narrower, pressed together, closed in, compact, and like sound in a smaller room. Some listeners mentioned that they also focused on timbre (sound color) and sound quality. A few listeners said they sometimes felt that there was too much echo, resonance, and reverberation. Most of these comments applied to the speech and mixed signals. None of the listeners reported any difficulty with speech intelligibility, though. One person said that more reflections made the music sound more alive and compared it to listening with the door wide open, as opposed to listening with the door almost shut. Another person said that it felt like being in a concert hall. SUMMARY The Music Widening algorithm was developed to improve the poor sound image that can result from listening to music that is streamed directly to hearing aids. A simple mirror-image model of a room is used to calculate the contributions of the direct sound and the room reflections. By replacing each contribution with a virtual sound source, it is possible to create the impression of sound sources in a simulated room. This sound image is externalized, i.e., perceived to be outside the head. This is in contrast to traditional stereo, where the sound image can be lateralized, but is still perceived as being inside the head. The results from the listening experiment indicate that applying this processing to a streamed music signal can dramatically improve the spatial impression.
Read more