All papers
Speech synthesisFull-paper digest

SELFIX: An Interactive System for Natural Self-Voice Approximation

Pavo Orepic, Steven Moran, Volker Dellwo

Code & resourcesosf.io/6nxzw

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

12.3 KB · Ready to paste

Preview copied content

TL;DR — SELFIX is a browser-based interactive system that uses perceptually motivated acoustic dimensions to model natural self-voice approximation, demonstrating that users consistently boost low-frequency energy by 6.68 dB while males additionally lower pitch and increase vocal-tract length.

Key contributions

  • Designed SELFIX, an interactive web-based framework integrating spectral filtering, source-filter theory, and psychoacoustic voice-quality parameters.
  • Conducted a proof-of-concept user study (N = 25) using randomized, unlabeled sliders to eliminate starting-position and centering biases.
  • Identified a universal requirement for low-frequency amplification (mean 6.68 dB boost at 600 Hz low-shelf) and dismissed a universal fixed trapezoid filter.
  • Uncovered systematic gender-specific adjustments in identity-related parameters, with males lowering pitch (-0.74 semitones) and increasing VTL (ratio 1.01).

Problem

People frequently experience discomfort and perceived unnaturalness when listening to recordings of their own voice due to the absence of bone-conduction filtering (internal skull/tissue vibrations) present during natural speech. Previous attempts relied on unguided spectral equalization or fixed transfer filters across vast acoustic spaces, failing to converge on universal solutions or map to perceptual dimensions. Overcoming this requires grounding voice transformation in perceptually relevant, low-dimensional acoustic features that capture both spectral shaping and voice identity.

Method

The SELFIX client runs in a standard web browser communicating via FastAPI and Uvicorn. The backend audio processing pipeline combines source-filter manipulations via Parselmouth and spectral filtering via SciPy. Incoming audio is converted to mono, analyzed for global pitch, and processed through: (1) pitch and vocal-tract-length (VTL) modifications using PSOLA resynthesis, (2) source-filter decomposition via linear predictive coding (LPC), (3) optional source envelope modifications (e.g., HNR or jitter, unused here), (4) parametric spectral filtering using low-shelf and trapezoid-shaped responses, and (5) peak-controlling and soft-clipping for stable playback.

All parameters are constrained to physiologically plausible ranges, with 35 features available though the proof-of-concept utilized four: F0 shift, VTL scaling, a 600 Hz low-frequency shelving filter (LS), and Vurma's trapezoid filter (preserving 1.7–3.2 kHz). Sliders were presented with randomized initial positions (20-80% range) and a neutral-reference shift of ±5-10% to prevent anchoring biases. In Part 1, users adjusted single randomized sliders over 20 trials, rating match and confidence via visual analog scales (0-100). In Part 2, users freely manipulated all four sliders simultaneously across 3 trials to examine multi-parameter interactions.

Experimental setup

Evaluated with 25 participants (13 male, 12 female) recording the utterance 'Hi, how are you?' on a Lenovo 21ML008YMZ laptop using integrated microphones and speakers. Metrics included final slider displacement relative to initial values, visual analog scale (VAS) match and confidence ratings, reaction times, number of interactions, and PCA/correlation variance structures. Implemented in Python using NumPy, SciPy, statsmodels, and scikit-learn.

Results

In Part 1, participants significantly boosted the 600 Hz low-shelf filter by an average of 6.68 dB (t(124) = 6.77, p < 0.001, d = 0.61), whereas the trapezoid filter showed no significant group effect (t(124) = 0.33, p = 0.741). Male participants significantly lowered pitch by -0.74 semitones (t(64) = -3.79, p < 0.001) and increased VTL to 1.01 (t(64) = 3.84, p < 0.001), while females showed a non-significant pitch reduction and no VTL shift. Subjective match and confidence ratings were high (mean match 76.46 ± 19.39, confidence 76.08 ± 18.90). In Part 2, simultaneous multi-slider manipulation yielded significantly higher match ratings than Part 1 (t(24) = 4.36, p < 0.001, d = 0.87) with high inter-trial consistency (pairwise r >= 0.75).

Parameter / ConditionUnitMean Adjustment (Overall / Male / Female)Statistical Significance
600 Hz Low-Shelf (LS)dB+6.68 dB (All)t(124) = 6.77, p < 0.001
Trapezoid Filter%0 (No effect)t(124) = 0.33, p = 0.741
Pitch (F0) Shiftsemitones-0.74 (Male) / -0.28 (Female)Male: t(64) = -3.79, p < 0.001
Vocal Tract Length (VTL)ratio1.01 (Male) / 1.00 (Female)Male: t(64) = 3.84, p < 0.001

Limitations

The study relies on a small sample size (N = 25) examining a single spoken utterance ('Hi, how are you?') captured exclusively through standard built-in laptop hardware in a non-studio environment. The evaluation is limited to four out of 35 possible SELFIX features, and gender-specific adjustments may intertwine perceptual familiarity with aspirational self-representation rather than pure acoustic matching.

Why read this

Researchers and engineers building voice cloning, real-time voice modification, or clinical speech tools should read this to understand how low-dimensional, perceptually constrained interfaces capture natural self-voice perception better than unguided equalizers.

Code

Applications

Perceptually calibrated AI voice cloning, real-time voice transformation tools, gender-affirming voice therapy support, and clinical interventions for auditory-verbal hallucinations or neuroprostheses.

Institutions

University of Zurich, University of Neuchâtel

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-1395