
Senior Voice AI Engineer / Voice Platform Engineer
Parse AI · Lebanon
Remote
About the job
Parse AI is hiring a Senior Voice AI Engineer / Voice Platform Engineer to help develop our scalable AI call center platform for hospitality operators.
Important: To be considered for this role, please complete this short screening form: https://forms.gle/HJAZqcLczzcXaH8H6
Only candidates who submit the form will be reviewed.
We are looking for a Senior Voice AI Engineer / Voice Platform Engineer to own the real-time voice layer of our platform. This includes telephony, streaming audio, STT, LLM orchestration, TTS, barge-in, latency, call transfers, multilingual voice, and production reliability.
This is not a chatbot role. This is a production voice engineering role for someone who understands how real phone calls behave.
Role Mission
Own the real-time voice pipeline from phone and SIP ingress through speech recognition, turn detection, LLM orchestration, text-to-speech, interruption handling, and transfer flows.
Your goal is to make AI phone calls feel natural, reliable, low-latency, and scalable across many restaurants, hotels, tenants, and call center environments.
What You Will Work On
Design and improve real-time STT → LLM → TTS voice pipelines
Tune voice providers such as Deepgram, OpenAI, ElevenLabs, Cartesia, Hamsa, or similar tools
Improve VAD, endpointing, silence handling, turn-taking, and barge-in behavior
Build pronunciation and recognition lexicons for names, restaurant terms, hotel terms, accents, and multilingual use cases
Work on Twilio Media Streams, SIP trunking, call transfers, warm transfers, and telephony reliability
Measure and improve voice latency across STT, LLM, TTS, network, and telephony layers
Debug production voice issues such as garbled audio, clipping, echo, slow responses, false interruptions, missed interruptions, and failed transfers
Help scale concurrent call capacity across multiple customers and environments
Partner with AI and backend engineers to make sure the agent experience is fast, safe, and production-ready
What We Are Looking For
We are looking for someone with real experience building and operating production voice systems.
Strong candidates will have experience with:
Real-time voice agents or voice platforms
STT, TTS, and speech-to-speech systems
Twilio Media Streams, SIP trunks, WebRTC, LiveKit, Telnyx, or similar systems
Voice activity detection, endpointing, barge-in, and interruption handling
Streaming audio, audio buffers, and telephony audio formats
Voice latency measurement and optimization
Production debugging of real phone call issues
Multilingual or accent-heavy voice environments
Provider evaluation and tuning across STT and TTS vendors
Strong Plus
Experience with restaurant, hotel, hospitality, or contact center voice agents
Experience with LiveKit Agents, Rapida, Modal, or serverless voice runtimes
Experience supporting 50+ concurrent calls
Experience with Arabic, French, Polsih, MENA accents, or multilingual speech systems
Experience building observability for voice quality, turn latency, interruptions, and transfers
Experience building voice AI evals, simulations, and pre-production testing frameworks for latency, accuracy, barge-in, multilingual performance, and call readiness
What This Role Is Not
This is not a general LLM research role.
This is not a frontend role.
This is not generic backend CRUD work.
This is not a simple voice API integration role.
We need someone who can help make the actual phone call work in production.
Ideal Candidate
You are likely a strong fit if you:
Think in milliseconds, audio formats, and real call behavior
Have shipped real voice systems used by actual callers
Can explain what broke in production and how you fixed it
Understand that voice quality depends on telephony, STT, TTS, latency, interruption handling, and reliability working together
Care deeply about making AI calls feel natural, fast, and dependable
Location and Work Style
This is a full-time remote role. Candidates should be able to overlap with MENA and Europe working hours, with occasional US timezone overlap when needed.
We are especially interested in candidates who have shipped real production voice systems and can bring practical experience from live deployments.
Ready to apply?Apply now