<?xml version="1.0" encoding="UTF-8"?><xml><records><record><source-app name="Biblio" version="7.x">Drupal-Biblio</source-app><ref-type>47</ref-type><contributors><authors><author><style face="normal" font="default" size="100%">You, Yuhuan</style></author><author><style face="normal" font="default" size="100%">Qian, Yufan</style></author><author><style face="normal" font="default" size="100%">Tianshu Qu</style></author><author><style face="normal" font="default" size="100%">Bin Wang</style></author><author><style face="normal" font="default" size="100%">Lv, Xueyang</style></author></authors></contributors><titles><title><style face="normal" font="default" size="100%">Flow-HOA: Generative Joint Optimization for Ambisonics Encoding via Flow Matching</style></title><secondary-title><style face="normal" font="default" size="100%">the AES 160th Convention</style></secondary-title></titles><dates><year><style  face="normal" font="default" size="100%">2026</style></year><pub-dates><date><style  face="normal" font="default" size="100%">28-30 May </style></date></pub-dates></dates><urls><web-urls><url><style face="normal" font="default" size="100%">https://aes.org/publications/elibrary-page/?id=23190</style></url></web-urls></urls><pub-location><style face="normal" font="default" size="100%">Copenhagen, Denmark</style></pub-location><pages><style face="normal" font="default" size="100%">10293</style></pages><language><style face="normal" font="default" size="100%">eng</style></language><abstract><style face="normal" font="default" size="100%">Higher-Order Ambisonics (HOA) encoding from sparse, irregular microphone arrays remains a critical challenge for consumer spatial audio capture in immersive communication and XR. We propose Flow-HOA, a generative framework that jointly optimizes a multi-dimensional objective encompassing time-domain, spectral, and spatial fidelity while producing a deployable, time-invariant bank of Finite Impulse Response (FIR) encoding filters. Using conditional flow matching, the model learns to map a simple prior distribution to the target distribution of FIR filtercoefficients. Training is guided by a composite loss that balances time-domain waveform fidelity, multi-resolution spectral consistency, sub-band energy preservation, and spatial directivity constraints. Objective evaluations on synthetically simulated data demonstrate improved performance over strong model-based baselines in both signal fidelity and spatial accuracy metrics. Subjective listening tests on real microphone array recordings further confirmthat Flow-HOA yields higher overall sound quality with reduced artifacts, demonstrating generalization from synthetic training data to real-world capture conditions.</style></abstract></record></records></xml>