Skip to content

fix: encode microphone recordings as valid PCM WAV - #3900

Open
HaokaiDing wants to merge 1 commit into
openai:mainfrom
HaokaiDing:fix/microphone-wav-sample-encoding
Open

HaokaiDing wants to merge 1 commit into
openai:mainfrom
HaokaiDing:fix/microphone-wav-sample-encoding

Conversation

@HaokaiDing

Copy link
Copy Markdown
Contributor
  • I understand that this repository is auto-generated and my pull request may not be merged

Changes being requested

Microphone.record() writes every capture dtype directly into an integer PCM WAV file. For float32, the IEEE float bytes are consequently interpreted as signed integers. For int8, silence (0) is interpreted as the most negative sample because 8-bit WAV PCM is unsigned.

Convert float32 recordings to 16-bit PCM and offset signed 8-bit recordings to unsigned PCM before writing the WAV, using the converted sample width. Preserve the existing int16, int32, and uint8 encodings and the original dtype/data when return_ndarray=True.

Additional context & links

Regression tests drive the public record() method through a mocked input device and decode the resulting WAV. They cover all five supported capture dtypes, stereo metadata, and ndarray returns. The float32 and int8 WAV cases fail before the fix; the other eight cases pass.

Validation on the latest main (3.16.1):

  • .venv/bin/python -m pytest -n 0 tests/lib/test_microphone.py tests/lib/test_audio.py -q: 14 passed.
  • Ruff lint and format checks on both changed files: passed.
  • Pyright on both changed files with the pinned repository tooling: 0 errors, 0 warnings.

Format references: sounddevice sample types, WAV PCM sample ranges.

@HaokaiDing
HaokaiDing requested a review from a team as a code owner September 18, 2026 19:02
Copilot AI lite review requested due to automatic review settings September 18, 2026 19:02

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@jnohclee-rgb

Copy link
Copy Markdown

AI-assisted independent offline byte-level corroboration at 044212e5efd47483ae68934004ee626aa0571c92 vs parent bccad31277a98f120bdce604a35a5774a0beaeae. Five fictional sample arrays through the actual _ndarray_to_wav helper, decoded back with Python wave and explicit little-endian PCM:

  • float32 [-2,-1,0,1,2]: parent writes four-byte float bit patterns labelled PCM; this head produces two-byte PCM [-32767,-32767,0,32767,32767], confirming clipping.
  • int8 [-128,0,127]: parent decodes as unsigned [128,0,127]; this head gives [0,128,255], correctly relocating silence.
  • uint8 and int16 controls: samples preserved.
  • float32 [NaN,+Inf,-Inf,0]: this head yields [0,32767,-32767,0] on NumPy2.5.3 and emits 'invalid value encountered in cast'. The NaN result should not be treated as a portable sanitization policy; it comes from NumPy's invalid cast.

The core float/int8 correction is corroborated. Please add an explicit non-finite input policy (reject, replace or documented handling) and a test, rather than relying on cast behavior; clipping alone leaves NaN unresolved. The out-of-range finite control is also useful beyond the existing normalized fixtures.

Scope: isolated helper and byte round-trip only; no recording, playback, sounddevice device access or native callback proof. Python3.14/Pydantic2.13.5; NumPy2.5.3 added to an isolated diagnostic directory, other dependencies shared. Archived sources are not a historical locked environment. Network denied during checks; no API or model activity.

Standalone reproducer
import json,wave,warnings,numpy as np
from openai.helpers.microphone import Microphone
rows=[]
for name,values,dtype in [('float_edges',[-2,-1,0,1,2],np.float32),('float_nonfinite',[np.nan,np.inf,-np.inf,0],np.float32),('signed8',[-128,0,127],np.int8),('unsigned8',[0,128,255],np.uint8),('signed16',[-32768,0,32767],np.int16)]:
 data=np.array(values,dtype=dtype);m=Microphone(channels=1,dtype=dtype)
 with warnings.catch_warnings(record=True) as w:
  warnings.simplefilter('always');name_,buffer,media=m._ndarray_to_wav(data)
 with wave.open(buffer,'rb') as reader:
  width=reader.getsampwidth();samples=np.frombuffer(reader.readframes(reader.getnframes()),dtype={1:'u1',2:'<i2',4:'<i4'}[width]).tolist();rows.append({'case':name,'sample_width':width,'decoded_pcm':samples,'warnings':[str(x.message) for x in w]})
print(json.dumps(rows))

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants