- Research Article
- 10.1016/j.csl.2025.101924
Audiovisual speech enhancement and voice activity detection using generative and regressive visual features
- Jul 01, 2026
- Computer Speech & Language
- Cheng Yu + 3 more +3
Publications from 2021 to 2026
Showing 10 of 97 papers
Audiovisual speech enhancement and voice activity detection using generative and regressive visual features
Real‐Time Neural Materials on Mobile VR
Abstract Virtual Reality (VR) applications aim to create an immersive virtual world, which demands a high level of visual realism. The analytical material models commonly used in VR often fall short of reproducing complex real‐world appearances. Recently, neural materials have emerged as a promising alternative, offering a compact yet effective representation of real‐world materials. Deploying neural materials on low‐power mobile VR devices poses significant challenges due to the computational complexity of neural networks and the high display resolution and frame rate requirements of VR devices (commonly 72+ frames per second). We address these challenges by leveraging texture‐space shading with spatiotemporal computation amortization, driven by a compact, coarse‐to‐fine neural material model of extremely low capacity. Thanks to our distillation training scheme, our compact neural materials achieve visual quality comparable to NeuMIP [KMX*21] at a much lower cost. Our method reaches over 90 FPS on a mobile VR device (Meta Quest 3) even under multiple light sources.
Read moreSubmillisecond-response LCD for low power field-sequential-color virtual reality displays.
We report an optimized transmissive fringe-field-switching (FFS) liquid crystal display (LCD) employing a low viscosity nematic mixture ZOC-5322. Submillisecond average gray-to-gray (GTG) response time is achieved at 40 °C by using a two-step overdrive and undershoot driving scheme. Such a fast response time enables field-sequential-color (FSC) operation to mitigate color breakup. By removing the lossy color filters, such an FFS LCD triples the resolution density and optical efficiency to fulfill 1-arcminute visual acuity and low power requirements of virtual reality (VR) headsets. Moreover, our optimized non-virtual-wall FFS cell exhibits a 1.64x faster average GTG response time and ∼ 6% higher transmittance than those of conventional virtual-wall counterparts.
Read moreInteractive design of developable surfaces by patch-based learning
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
Shima Imani, Seungwhan Moon, Adel Ahmadyan, Lu Zhang, Ahmed Kirmani, Babak Damavandi. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track). 2026.
Read moreA Nonconforming Formulation of Cloth
State-of-the-art cloth simulations rely on linear triangular elements in mass-spring or continuum based finite element formulations. These methods typically decompose the surface energy density into in-plane (shearing and stretching) and out-of-plane (bending) components, with bending energies modeled using discrete mean curvature measures. While effective, they are prone to mesh-dependent behavior and locking. Higher-order formulations can mitigate these issues, but their adoption poses significant challenges due to the requirement for continuity of basis functions’ derivatives across element boundaries to accurately represent surface curvature. We introduce a novel continuum-based approach that addresses the limitations of existing methods without requiring globally smooth (H2-continuous) basis functions. Our method uses non-conforming function spaces and weakly enforces the continuity of tangent basis through carefully derived interface terms. In fact, the proposed method builds on Interior Penalty methods, which we adapt to effectively handle simulations of curved surfaces. Our approach uses standard Lagrangian basis functions, and supports straightforward extension to high-order bases, while adhering to the in-plane/out-of-plane decoupling paradigm widely adopted in cloth simulation. We demonstrate the robustness and versatility of our method through garment simulations, illustrating its ability to handle complex deformations and a variety of bending behaviors with high fidelity.
Read moreComprehensive characterization of human color discrimination thresholds
Color discrimination thresholds—the smallest detectable color differences—provide a benchmark for models of color vision, enable quantitative evaluation of eye diseases, and inform the design of display technologies. Despite their importance, a comprehensive characterization of these thresholds has long been considered intractable due to the psychophysical curse of dimensionality. Here, we address this challenge using a novel semi-parametric Wishart Process Psychophysical Model (WPPM), which leverages the feature that the internal noise limiting color discrimination varies smoothly across stimulus space. The model was fit to data collected with a non-parametric adaptive trial-placement procedure, enabling efficient stimulus selection. Together, through the combination of adaptive trial placement and post hoc WPPM fitting, we achieved comprehensive characterization of color discrimination in the isoluminant plane with only ~6,000 trials per participant (N = 8). Once fit, the WPPM allows readouts of discrimination performance for any stimulus pair. We validated these readouts against 25 probe psychometric functions, measured with an additional 6,000 trials per participant held out from model fitting. In conclusion, our study provides a foundational dataset for color vision, and our approach generalizes beyond color to any domain in which the internal noise limiting performance varies smoothly across stimulus space, offering a powerful and efficient method for comprehensively characterizing various perceptual discrimination thresholds.
Read moreIndividual Differences in Training Naive Listeners to Localize Spatial Audio in Virtual Reality
It has been widely believed that a key factor in creating realistic spatial audio in virtual reality (VR) is the head-related transfer function (HRTF), which is unique to each individual, but costly to measure for widespread use. This study investigates the effects of HRTF personalization and training on sound localization accuracy in VR. Two experiments were conducted: Experiment 1 compared naive listeners and those who underwent brief training on localization tasks using personalized versus generic HRTFs; Experiment 2 used a within-subject design to assess training effects over two sessions. Results show that accurately localizing sound can be a difficult task for many participants the first time. Training significantly improves localization accuracy, reducing errors and confusions, and enabling many initially non-sensitive listeners to perceive spatial audio effectively. Although HRTF personalization yielded a statistically significant benefit, the effect was small, primarily improving elevation perception at extreme angles. These findings suggest that generic HRTFs combined with user training may suffice for most VR applications.
Read moreNovel Diffusion Models for Multimodal 3D Hand Trajectory Prediction
A Unified Framework for Evaluating DNN-Based Feedforward, Feedback, and Hybrid Active Noise Cancellation
Deep neural network (DNN)-based acoustic noise cancellation excels at modeling complex non-linear relationships in signal patterns that are difficult for linear filters to handle. Recent work has studied DNN-based feedforward (FF) and feedback (FB) control structures. However, their inconsistent experimental settings prevent fair comparison, and limited evaluation in a narrow range of acoustic conditions hinders a comprehensive understanding of their effectiveness. In this work, we present a unified framework that realizes FF and FB control structures to better understand their cancellation mechanisms under various acoustic conditions. In addition, it enables a hybrid (HB) control structure that combines the FF and FB approaches, which has not been evaluated against standalone DNN-based FF and FB configurations. We perform systematic evaluations in various settings, including different disturbing signals, reverberation conditions, source positions, room sizes, and mismatched secondary paths. The evaluation results show that the effectiveness of FF depends strongly on the modeling of the primary path and the acoustic environment, the performance of FB varies less under different acoustic conditions, and HB integrates the advantages of FF and FB.
Read more