G1 Audio Playback & Recording — Working Configuration
======================================================
Date: 2026-04-07  (updated 2026-04-23)

UPDATE 2026-04-23
------------------
Three scripts previously used the broken chunked-PlayStream path. Fixed:

    g1_voice_recorder.py  —  replay path now sends ALL audio in ONE
                             _CallRequestWithParamAndBin call with a unique
                             stream_id. No more PlayStream.
    g1_play_wav.py        —  same fix; removed TtsMaker warm-up (not needed).
    g1_interactive_player.py  —  removed chunked loop (same stream_id across
                             chunks was silently dropping after the first).

Added three new scripts:

    g1_audio_devices.py      —  pactl wrapper: list sources/sinks, find by
                                keyword, auto-detect Hollyland mic, unmute,
                                set volume, parec quick-test.
    g1_edge_tts.py           —  edge-tts bilingual (AR+EN) TTS pipeline —
                                text → MP3 → 16 kHz WAV → G1 speaker.
                                Voices default to ar-AE-HamdanNeural / en-US-GuyNeural.
    template_audio_script.py —  copy-paste starter demonstrating the
                                reliable primitives (prep mic, record via
                                parec, play in ONE call, always STOP).

Quick reference — the ONE audio play path that works on G1:
    from unitree_sdk2py.g1.audio.g1_audio_api import (
        ROBOT_API_ID_AUDIO_START_PLAY, ROBOT_API_ID_AUDIO_STOP_PLAY)
    c._Call(ROBOT_API_ID_AUDIO_STOP_PLAY, json.dumps({"app_name": APP}))
    time.sleep(0.3)
    sid = f"s_{int(time.time()*1000)}"
    param = json.dumps({"app_name": APP, "stream_id": sid,
                        "sample_rate": 16000, "channels": 1, "bits_per_sample": 16})
    c._CallRequestWithParamAndBin(ROBOT_API_ID_AUDIO_START_PLAY, param, list(pcm))
    time.sleep(duration + 0.5)
    c._Call(ROBOT_API_ID_AUDIO_STOP_PLAY, json.dumps({"app_name": APP}))

------------------



BUILT-IN SPEAKER PLAYBACK (via Unitree Voice Service)
-----------------------------------------------------
Method: _CallRequestWithParamAndBin (NOT PlayStream)
Key discovery: JSON params MUST include sample_rate, channels, bits_per_sample
Required: Send ALL audio in ONE call (streaming chunks does NOT work)
Required: Unique stream_id per play (timestamp-based), or audio won't repeat
Required: Call AUDIO_STOP_PLAY before each new play to reset stream

Working code pattern:
    from unitree_sdk2py.g1.audio.g1_audio_api import *

    c._Call(ROBOT_API_ID_AUDIO_STOP_PLAY, json.dumps({"app_name": "p"}))
    time.sleep(0.3)

    sid = f"s_{int(time.time() * 1000)}"
    param = json.dumps({
        "app_name": "p",
        "stream_id": sid,           # MUST be unique each time
        "sample_rate": 16000,       # MUST include these format params
        "channels": 1,
        "bits_per_sample": 16,
    })
    c._CallRequestWithParamAndBin(ROBOT_API_ID_AUDIO_START_PLAY, param, list(pcm))
    time.sleep(duration + 1)
    c._Call(ROBOT_API_ID_AUDIO_STOP_PLAY, json.dumps({"app_name": "p"}))

Audio format: 16kHz, mono, 16-bit signed little-endian PCM
SDK required: Robot's gemini conda env (has _CallRequestWithParamAndBin)
Workstation SDK (g1_env): Does NOT have _CallRequestWithParamAndBin — must run on robot

What does NOT work:
    - PlayStream() method (sends data but no audio output)
    - Streaming chunks in a loop (only first call sometimes plays)
    - _CallBinary / _Call with PCM data (error code 100 from workstation SDK)
    - Running playback from .py script files (only python3 -c inline works reliably)
    - aplay / paplay through PulseAudio (Jetson platform-sound has no physical speaker)

TTS (Text-to-Speech):
    - English: works
    - Chinese: works
    - Arabic: does NOT work (falls back to Chinese)
    - speaker_id: 0, 1, 2 all work (different voices)

API IDs:
    ROBOT_API_ID_AUDIO_TTS         = 1001
    ROBOT_API_ID_AUDIO_ASR         = 1002  (returns 3104 — not available)
    ROBOT_API_ID_AUDIO_START_PLAY  = 1003
    ROBOT_API_ID_AUDIO_STOP_PLAY   = 1004
    ROBOT_API_ID_AUDIO_GET_VOLUME  = 1005
    ROBOT_API_ID_AUDIO_SET_VOLUME  = 1006
    ROBOT_API_ID_AUDIO_SET_RGB_LED = 1010
    APIs 1007-1009, 1011-1012: not registered (code 3103)

WIRELESS MICROPHONE RECORDING
------------------------------
Device: Hollyland Wireless Microphone (USB)
    Vendor: Shenzhen Hollyland Technology Co.,Ltd
    Product ID: 0007
    ALSA: card 2, device 0 (hw:2,0)
    PulseAudio source index: 3
    Source name: alsa_input.usb-Shenzhen_Hollyland_Technology_Co._Ltd_Wireless_microphone_C63X223T6MX-01.analog-stereo
    Format: s24le, 2ch, 48000Hz

CRITICAL: Mic was MUTED by default. Must unmute before recording.

Step-by-step mic setup commands (run on robot):
    # 1. View all PulseAudio sources (find mic index)
    pactl list sources short

    # 2. Check mic details (mute status, volume, format)
    pactl list sources | grep -A 10 "Wireless"

    # 3. Set wireless mic as default input source
    pactl set-default-source alsa_input.usb-Shenzhen_Hollyland_Technology_Co._Ltd_Wireless_microphone_C63X223T6MX-01.analog-stereo

    # 4. Unmute the mic (index 3)
    pactl set-source-mute 3 0

    # 5. Set mic volume to 100%
    pactl set-source-volume 3 100%

    # 6. Verify mic is capturing audio
    timeout 2 parec -d 3 --format=s16le --rate=16000 --channels=1 --raw > /tmp/raw_test.pcm
    python3 -c "
    import numpy as np
    a = np.fromfile('/tmp/raw_test.pcm', dtype=np.int16)
    print(f'Samples={len(a)}, min={a.min()}, max={a.max()}, std={a.std():.0f}')
    if a.std() > 50: print('MIC WORKS!')
    else: print('Still silent - check transmitter is ON')
    "

Recording method (parec, not PyAudio — PyAudio gives silence through pulse device):
    subprocess.Popen(['parec', '-d', '3', '--format=s16le', '--rate=16000', '--channels=1', '--raw'], stdout=subprocess.PIPE)
    time.sleep(duration)
    proc.terminate()
    raw = proc.stdout.read()

PyAudio recording through device index 25 (pulse) gives ALL ZEROS even after unmuting.
Must use parec with source index 3 directly.

FULL RECORD + PLAYBACK EXAMPLE (run on robot in gemini conda env):
    python3 -c "
    import time, wave, json, numpy as np, subprocess
    from unitree_sdk2py.core.channel import ChannelFactoryInitialize
    from unitree_sdk2py.g1.audio.g1_audio_client import AudioClient
    from unitree_sdk2py.g1.audio.g1_audio_api import *

    # Record 5 seconds via parec (wireless mic, source index 3)
    print('Recording 5 seconds... speak now!')
    proc = subprocess.Popen(
        ['parec', '-d', '3', '--format=s16le', '--rate=16000', '--channels=1', '--raw'],
        stdout=subprocess.PIPE)
    time.sleep(5)
    proc.terminate()
    raw = proc.stdout.read()
    audio = np.frombuffer(raw, dtype=np.int16)
    print(f'Recorded {len(audio)} samples, std={audio.std():.0f}')

    # Save to WAV
    path = '/home/unitree/SanadVoice/recorded voices/my_recording.wav'
    wf = wave.open(path, 'wb')
    wf.setnchannels(1)
    wf.setsampwidth(2)
    wf.setframerate(16000)
    wf.writeframes(audio.tobytes())
    wf.close()
    print(f'Saved: {path}')

    # Play back on G1 built-in speaker
    ChannelFactoryInitialize(0, 'eth0')
    c = AudioClient()
    c.SetTimeout(10.0)
    c.Init()
    c.SetVolume(100)
    c._Call(ROBOT_API_ID_AUDIO_STOP_PLAY, json.dumps({'app_name':'p'}))
    time.sleep(0.3)

    pcm = audio.tobytes()
    sid = f's_{int(time.time()*1000)}'
    param = json.dumps({
        'app_name': 'p',
        'stream_id': sid,
        'sample_rate': 16000,
        'channels': 1,
        'bits_per_sample': 16
    })
    c._CallRequestWithParamAndBin(ROBOT_API_ID_AUDIO_START_PLAY, param, list(pcm))
    time.sleep(len(audio)/16000 + 1)
    c._Call(ROBOT_API_ID_AUDIO_STOP_PLAY, json.dumps({'app_name':'p'}))
    print('Done')
    "

NETWORK & SSH
--------------
Robot IP: 192.168.123.164
SSH user: unitree
SSH pass: 123
SSH key: configured (ssh-copy-id done)
DDS interface on robot: eth0
DDS interface on workstation: enp3s0

Robot conda envs:
    - base: no unitree_sdk2py
    - gemini: has unitree_sdk2py with PlayStream + _CallRequestWithParamAndBin
    Python: /home/unitree/miniconda3/envs/gemini/bin/python3

RUNNING SERVICES ON ROBOT
---------------------------
PulseAudio: PID 1752 (holds hw:1,0 speaker and hw:2,0 mic)
sanad_webserver.py: PID 1815 (gemini_voice_v2)
sanad_voice.py: PID ~4640+ (gemini_voice)
Systemd: sanad_voice.service (auto-start)

PulseAudio sink (speaker): alsa_output.platform-sound.analog-stereo (card 1 APE device 0)
    - Volume: 100%, not muted
    - Does NOT produce audible output (no physical speaker on this path)
    - Built-in speaker is only accessible through Unitree voice RPC service

SanadVoice project speaker: Anker PowerConf USB (external, not always connected)
    Sink: alsa_output.usb-Anker_PowerConf_A3321-DEV-SN1-01.analog-stereo
    Source: alsa_input.usb-Anker_PowerConf_A3321-DEV-SN1-01.mono-fallback

RECORDED VOICES LOCATION
--------------------------
Robot: /home/unitree/SanadVoice/recorded voices/  (54 WAV files, English)
Robot: /home/unitree/SanadVoice/Audio/             (additional audio)
Workstation: ~/Robotics_workspace/yslootahtech/G1_Lootah/Audio_Recorder/DataG1/

MARCUS TTS CONFIGURATION (2026-04-08)
======================================
Backend: edge-tts (online, Microsoft Edge TTS API, no API key needed)
Requires: Internet connection on robot

Voices:
    Arabic:  ar-AE-HamdanNeural  (UAE male)
    English: en-US-GuyNeural     (US male)
    Note: For same voice both languages, use ar-AE-HamdanNeural for both
          (English will have slight Arabic accent)

Config: /home/unitree/Marcus/Config/config_Voice.json
    "en_backend": "edge_tts"
    "ar_backend": "edge_tts"
    "edge_voice_ar": "ar-AE-HamdanNeural"
    "edge_voice_en": "en-US-GuyNeural"

Pipeline:
    audio_api.speak(text, lang)
        → edge-tts generates MP3 (1-2s)
        → pydub converts MP3 → 16kHz WAV
        → _CallRequestWithParamAndBin → G1 speaker
        → mic muted during playback, unmuted after

Fallback: Built-in TtsMaker for English if edge-tts fails (no internet)

Mic: Auto-detected by scanning PulseAudio for "wireless"/"hollyland"/"usb"
    Auto-unmutes, sets volume 100%, sets as default source
    Uses parec (not PyAudio) for recording
    Muted during TTS playback to prevent self-listening

Files updated:
    /home/unitree/Marcus/API/audio_api.py     — XTTS + edge-tts + builtin backends
    /home/unitree/Marcus/Voice/marcus_voice.py — Uses audio_api's auto-detected mic
    /home/unitree/Marcus/Voice/xtts_server.py  — XTTS-v2 persistent server (for offline use later)
    /home/unitree/Marcus/Config/config_Voice.json — TTS config

Conda envs:
    marcus     (Python 3.8) — Brain, vision, controller, audio playback SDK
    marcus_tts (Python 3.10) — XTTS-v2, SILMA TTS (for future offline use)
    gemini     (Python 3.11) — Has _CallRequestWithParamAndBin

TTS OPTIONS TESTED
-------------------
| Backend          | Arabic | English | Offline | Speed  | Status        |
|------------------|--------|---------|---------|--------|---------------|
| edge-tts         | Yes    | Yes     | No      | 1-2s   | WORKING ✓     |
| Built-in TtsMaker| No     | Yes     | Yes     | 0s     | WORKING ✓     |
| XTTS-v2 (CPU)   | Yes    | Yes     | Yes     | ~14s   | WORKING (slow)|
| XTTS-v2 (GPU)   | Yes    | Yes     | Yes     | ~2s    | NO (CUDA 11.4 no cp310 wheel) |
| SILMA TTS        | Yes    | Yes     | Yes     | ?      | BROKEN (deps) |
| Piper TTS        | Yes    | No      | Yes     | ~1s    | BROKEN (piper_phonemize on ARM64) |
| gTTS             | Yes    | Yes     | No      | 1-2s   | WORKING (female only) |
| SpeechT5 local   | Yes    | No      | Yes     | ~2-3s  | NOT TESTED (model on robot) |

JETSON ORIN NX LIMITATIONS
---------------------------
- JetPack 5.1.1, L4T R35.3.1, CUDA 11.4
- NVIDIA only provides PyTorch cp38 wheels (Python 3.8)
- No cp310 GPU PyTorch available — must build from source
- scikit-learn TLS bug on aarch64 (libgomp cannot allocate memory in static TLS block)
  Fix: LD_PRELOAD=$(find env/lib -name "libgomp*.so*" | head -1)
- apt is broken (unmet dependencies), cannot install espeak-ng

SCRIPTS DEPLOYED ON ROBOT
---------------------------
/home/unitree/g1_interactive_player.py  — Interactive WAV selector + playback
/home/unitree/g1_play_wav.py            — CLI WAV player
/home/unitree/g1_tts_test.py            — TTS language test
/home/unitree/Marcus/Voice/xtts_server.py — XTTS-v2 persistent server

FILE INVENTORY (workstation /G1_Lootah/Audio_Recorder)
-------------------------------------------------------
Recording + replay:
    g1_voice_recorder.py        — workstation PyAudio record → scp → G1 replay
    g1_play_wav.py              — (on-robot) play any WAV on G1 speaker
    g1_interactive_player.py    — (on-robot) interactive WAV picker
    g1_voice_deploy.py          — deploy recordings to robot
Device management:
    g1_audio_devices.py         — pactl wrapper: list/find/auto-hollyland/quick-test
TTS:
    g1_edge_tts.py              — Microsoft Edge TTS bilingual (AR+EN)
    g1_tts_arabic.py            — Piper / tts_arabic offline TTS (ARM64 broken per above)
    g1_tts_test.py              — TTS smoke test (English only)
Templates + tests:
    template_audio_script.py    — copy-paste starter with the reliable primitives
    test_g1_speaker.py          — speaker init sanity check
    test_g1_pcm.py              — raw PCM playback test
    test_marcus_audio.py        — Marcus env audio pipeline end-to-end
Data:
    DataG1/                     — local recordings (WAV, 16 kHz mono int16)

Run recipes:
    # Workstation side
    python3 g1_voice_recorder.py record --name hello --seconds 5
    python3 g1_voice_recorder.py replay --name hello --ip 192.168.123.164
    python3 g1_edge_tts.py "Good morning" --lang en --save out.wav --no-play

    # On robot (gemini env)
    python3 g1_audio_devices.py auto-hollyland
    python3 g1_play_wav.py /path/to/file.wav
    python3 g1_edge_tts.py "مرحبا بكم" --lang ar
    python3 template_audio_script.py echo --seconds 3
