The four-stage pipeline
Listen
GirGit AI captures interview audio from active video calls and converts spoken questions into text in real time. This stage is the foundation of how the assistant works during live conversations.
Understand
Questions are categorised by intent — technical, behavioural or coding. The platform combines this with resume details and job requirements to frame the right response.
Generate
A language model processes the interview context and creates relevant responses aligned with the question type, professional experience and role-specific expectations.
Display
Responses are delivered through a private overlay interface within seconds, providing immediate interview guidance.
How the assistant captures system audio
The most common question about how these tools work is where the audio comes from. GirGit AI is a desktop application, so it captures the interview audio at the operating-system level — the same sound your call is already playing through your speakers or headphones — rather than reading anything inside Zoom, Teams or Meet.
This matters for two reasons. First, it means the interviewer's voice is transcribed straight from your system's audio output, so nothing has to be pasted or triggered manually. Second, because capture happens on your own machine and outside the meeting app, the video platform has no awareness of it — the assistant never joins the call as a participant or bot.
- Interviewer audio is captured from your system output (the call audio you already hear) and transcribed in real time.
- Your own microphone is untouched — the assistant reads the incoming question, it does not speak or inject audio into the call.
- No meeting bot joins the call; capture is local to your device, which is why it works identically on Zoom, Teams and Meet.
Because audio is captured and processed locally during the session, the interviewer sees a normal call — no extra participant, no recording notice from the assistant.
Real-time transcription with no delay or trigger
Interview conversations are transcribed continuously as participants speak, removing delays between question detection and response generation. The transcription system is built to handle varied accents, speaking speeds and call-quality conditions.
Context-aware answer generation
Responses are generated using the interview question, your resume and the job description together, rather than relying on isolated prompts. This aligns guidance with your professional experience and the role's requirements.
The invisible overlay
Generated responses appear through a private desktop overlay designed to stay outside the visibility of screen-sharing and screen-capture tools. Additional controls keep it separate from standard task-switching interfaces and desktop navigation.
Privacy and processing — interview audio is processed locally during active sessions. Transcripts and generated responses are not stored on GirGit AI servers, helping you keep control of interview-related information.
