This system manual serves as the single authoritative operational, mathematical, architectural, and developmental specification for the Sky Pro AI Calling Agent autonomous voice platform. It covers every component from network-level VoIP/SIP ingest, digital signal processing (DSP), real-time large language model inference, and client-side spreadsheet parsing, to full campaign analytics and automated callback queuing. All workflows described herein reflect the active production build.
Global tele-calling operations and customer contact centers are encountering systemic operational inefficiencies that erode corporate operating margins, compromise lead acquisition quality, and cause customer dissatisfaction.
Modern outbound tele-marketing and tele-sales teams operate under severe psychological strain. Human agents are expected to place 150 to 250 manual calls daily. Of these, less than 20% connect to live prospects, and fewer than 5% express qualified buying intent. This creates rapid agent burnout, resulting in an annualized turnover rate exceeding 45% across international business process outsourcing (BPO) centers.
Traditional DTMF IVR systems are universally detested by consumers. Callers are forced through lengthy recorded menus before reaching a representative, resulting in average drop-off rates of 32%. Furthermore, keypad IVRs are incapable of understanding nuance, interpreting urgent problems, or cross-referencing caller history in real time.
When a prospect says 'call me back at 6 PM', in 70% of legacy operations that lead is lost. Human agents take unstructured notes on paper or fail to set reminders. Without automated parsing, callbacks slip through the cracks, wasting valuable customer acquisition investments.
Sky Pro AI Calling Agent is an enterprise-grade autonomous telephony platform engineered from the ground up to deliver bidirectional, real-time conversational voice experiences at scale. It acts as an automated digital workforce, conducting complex conversations with sub-800 millisecond response times.
Sky Pro AI Calling Agent transforms enterprise customer communications by replacing rigid keypad trees with natural conversations. Human staff are freed to focus on high-value escalations, closing deals, and managing strategic client relationships.
Every feature maps to a specific operational failure observed in traditional tele-calling environments.
| Operational Failure | Traditional BPO Approach | Sky Pro Solution |
|---|---|---|
| Inconsistent Pitch Quality | Tired agents skip key qualification points; pitches vary wildly. | Deterministic prompt directives ensure consistent delivery on every call. |
| Unstructured Callback Losses | Manual notes on spreadsheets; callbacks frequently forgotten. | NLP temporal extraction automatically captures commitments into the Callback Scheduler. |
| Speech Collisions & Talk-over | High latency causes machine and human to speak simultaneously. | Client-side AudioWorklet barge-in instantly purges output audio buffers. |
| Lead Upload Bottlenecks | IT teams manually run SQL scripts to populate dialer lists. | SheetJS (excelParser.ts) allows drag-and-drop spreadsheet ingestion directly in browser. |
| Lack of Granular Telemetry | Call records only log connection time; outcomes unanalyzed. | Turn-by-turn dialogue logging, latency breakdowns, and auto disposition tagging. |
Sky Pro AI Calling Agent enforces role-based access control across three primary user tiers.
Scope: Telephony infrastructure, SIP gateways, LLM system directives, and server maintenance.
Permissions: Full, unrestricted access to all backend routes, logs, database migrations, and operational settings.
Scope: Outbound dialer operations, lead ingestion, schedule pacing, and conversion tracking.
Permissions: Access to Metrics, Dialer Queue, Scheduler, Calls, and Reports. Restricted from DSP configurations.
Scope: Dialogue verification, compliance tracking, and conversational auditing.
Permissions: Read-only access to Call History, Conversation Modals, Reports, and analytical exports.
A decoupled, high-performance client-server topology designed for low-latency streaming and high concurrent throughput.
| Subsystem | Technology / Library | Operational Functionality |
|---|---|---|
| Frontend Runtime | React 18 + Vite + TypeScript | SPA providing fast UI rendering and state handling. |
| UI & Components | Tailwind CSS + Lucide Icons | Responsive, modern interface styling for dashboards. |
| File Ingestion | SheetJS (xlsx) | Client-side Excel parsing transforming spreadsheets into dialer queues. |
| Backend Core | FastAPI (Python 3.11) | Asynchronous REST endpoints and WebSocket stream controllers. |
| AI Reasoning | Google Gemini 1.5 Pro / Flash | Context-aware dialogue understanding and response generation. |
| Voice Pipeline | Streaming STT & Neural TTS | Real-time bidirectional speech-to-text and text-to-speech. |
| Database ORM | SQLAlchemy + Alembic | Relational mapping, query execution, and migration tracking. |
| Web Server Proxy | Nginx (Alpine) | High-performance reverse proxy routing HTTP and WebSocket traffic. |
| Containerization | Docker + Docker Compose | Isolated multi-container production environments. |
Clean architectural layout separating backend voice pipelines, frontend user interfaces, and production deployment scripts.
Explore the 6 core functional modules of the Sky Pro AI Calling Agent ERP system.
Operational Purpose: The Metrics Dashboard is the central command cockpit for contact center supervisors providing real-time visibility into call volumes, lead conversions, and engagement.
Operational Purpose: The Dialer Queue eliminates manual database imports by enabling campaign managers to drag and drop contact spreadsheets directly into the browser via excelParser.ts.
Operational Purpose: When prospects state they are unavailable, the Callback Scheduler captures the follow-up commitment. Gemini detects the intent, extracts the temporal target, and adds the contact to the queue.
Operational Purpose: The Calls Ledger provides a detailed record of every completed voice interaction. Supervisors can review audio recordings, evaluate qualification accuracy, and inspect outcomes.
Operational Purpose: The Softphone Keypad allows voice engineers to initiate manual test calls, verify speech pipeline configurations, and connect live agents to ad-hoc numbers.
Operational Purpose: The Reports module is the platform's BI suite, allowing campaign managers to filter data across multiple parameters and export datasets for external CRM integration.
The inbound pipeline replaces rigid keypad menus with natural, open-ended conversational routing.
Customer dials in; SIP gateway routes the call via WebSocket to the FastAPI backend.
Agent introduces itself using inbound_prompt.txt: 'Hello! How can I help you today?'
Caller's voice is processed into streaming audio chunks and STT converts audio to text.
Gemini LLM parses the transcript, matching the request against knowledge bases.
Neural TTS synthesizes the response, streaming audio back with sub-800ms latency.
Conversation is summarized, classified, and saved to call_logs.
The outbound pipeline automates lead qualification and sales outreach at scale.
Campaign manager uploads an Excel lead list in the Dialer Queue interface.
Clicking Start triggers the outbound dialer up to configured concurrency limits.
VAD distinguishes between live human speech and voicemail systems.
Agent initiates using outbound_prompt.txt, confirming identity and presenting the pitch.
LLM addresses questions regarding pricing, features, and timing.
If interested, agent secures a demo slot. If busy, schedules a follow-up callback.
Call assigned disposition tag (Hot, Warm, Cold, Demo Booked) and written to database.
Sky Pro uses digital signal processing (DSP) to filter background noise and maintain audio quality.
Continuously analyzes incoming audio frames, distinguishing human speech from background silence.
Prevents the caller's audio from looping back into the microphone when using speakerphones.
Attenuates steady-state background noises such as air conditioning hum or street traffic.
Instantly clears audio buffers when the user speaks, ensuring smooth conversational flow.
Sky Pro targets an end-to-end response time under 800 milliseconds.
| Pipeline Stage | Target Latency | Optimization Method |
|---|---|---|
| VAD Speech Detection | 50 - 100 ms | Rolling energy thresholds to detect speech start and end. |
| Streaming STT | 150 - 200 ms | Transcribes audio chunks progressively over WebSockets. |
| LLM First Token Generation | 200 - 250 ms | Streams tokens from Gemini as generated, not waiting for complete responses. |
| Neural TTS Synthesis | 100 - 150 ms | Synthesizes audio from initial token chunk while subsequent text generates. |
| Network Buffer & Ingress | 50 - 100 ms | Compact PCM audio packets and optimized WebSocket routing. |
| Total End-to-End Latency | 550 - 800 ms | Natural conversational pacing without awkward pauses. |
Establishing behavioral boundaries and conversational structures for both inbound and outbound campaigns.
The file backend/Scripts/inbound_prompt.txt defines the system prompt for incoming customer service calls.
Identity & Persona: Define the agent as professional, patient, and knowledgeable.
Call Opening: Standard greeting acknowledging the customer.
Knowledge Retrieval: Instructions for querying internal documentation.
Escalation Rules: Clear guidelines for transferring callers to human supervisors.
Polite Wrap-Up: Standard closing verifying all questions are addressed.
The file backend/Scripts/outbound_prompt.txt establishes guidelines for tele-sales campaigns.
Contact Confirmation: Verify prospect's identity before delivering the pitch.
Value Proposition: Present a concise 15-second summary of the offer.
Objection Handling: Structured approaches to pricing concerns and timing issues.
Demo Conversion: Guide interested prospects toward scheduling a demonstration.
Callback Capture: If unavailable, capture preferred callback time for the scheduler.
Database models implemented using SQLAlchemy in backend/database.py.
| Table Name | Primary / Foreign Keys | Attributes & Types |
|---|---|---|
| call_logs | id UUID PK | session_id (VARCHAR UNIQUE), direction (INCOMING/OUTGOING), phone_number, duration_seconds (INT), lead_status (HOT/WARM/COLD), demo_booked (BOOL), recording_url, created_at (DATETIME). |
| callbacks | id UUID PK, call_log_id UUID FK | phone_number, requested_time (DATETIME), status (PENDING/COMPLETED/CANCELLED), notes (TEXT), assigned_agent_id. |
| queue_contacts | id UUID PK | contact_name, phone_line, dial_status (QUEUED/DIALING/COMPLETED/FAILED), retry_attempts (INT), outcome_tag, created_at. |
| transcripts | id UUID PK, call_log_id UUID FK | speaker_role (USER/AGENT), content_text (TEXT), turn_latency_ms (INT), sentiment_score (FLOAT), timestamp. |
REST endpoints for state operations; WebSockets for real-time audio streams.
| Endpoint / Route | Method / Protocol | Operational Function |
|---|---|---|
| /api/metrics | GET | Fetches aggregated call center metrics. |
| /api/queue/contacts | GET, POST | Retrieves or batch-inserts contact records. |
| /api/queue/start | POST | Starts outbound pacing dialer execution. |
| /api/scheduler/callbacks | GET, PATCH | Lists pending callbacks and updates statuses. |
| /api/calls/history | GET | Retrieves paginated call logs with recording links. |
| /api/reports/export | GET | Exports filtered call analytics as CSV download. |
| /ws/voice/stream | WebSocket | Full-duplex audio stream for real-time STT and voice playback. |
Step-by-step guidelines for administrators, campaign managers, and quality analysts.
SOP 1: System Health Verification
docker-compose logs -f backend.SOP 2: Prompt Configuration & Updates
backend/Scripts/ to inspect active prompt files.SOP 1: Ingesting Lead Lists & Starting Campaigns
SOP 2: Resolving Callbacks
SOP 1: Auditing Turn-by-Turn Transcripts
SOP 2: Reviewing Latency Telemetry
Sky Pro AI Calling Agent runs as containerized microservices managed via Docker Compose and Nginx.
| Command Syntax | Operational Function |
|---|---|
| docker-compose -f deploy/docker-compose.yml up --build -d | Builds and launches containers in detached mode. |
| docker-compose -f deploy/docker-compose.yml ps | Verifies the operational status of all running containers. |
| docker-compose -f deploy/docker-compose.yml logs -f backend | Streams real-time backend API and voice processing logs. |
Procedures for regular database backups, disaster recovery, and data migration.
Generate compressed database backups:
docker-compose down.alembic upgrade head.| Symptom | Probable Cause | Remediation Procedure |
|---|---|---|
| Turn-taking latency exceeds 1500ms | Slow token generation or network packet delay. | Switch to lower-latency Gemini tier; check regional network latency. |
| Agent cuts off callers prematurely | VAD silence threshold configured too short. | Increase silence detection threshold in voice.py from 400ms to 700ms. |
| Spreadsheet fails to import | Invalid or missing column headers in Excel file. | Ensure headers: Name, Phone, and Company are present. |
| No audio playback in browser | Browser autoplay policy blocking WebAudio contexts. | Click anywhere in the dashboard to unlock the WebAudio context. |
Common questions regarding Sky Pro AI Calling Agent architecture, capabilities, and configurations.
A: Yes. The platform includes active barge-in detection. When user speech is detected, the client immediately signals the backend, clearing outbound audio buffers so the user can speak naturally.
A: On inbound calls, the system reads caller ID from SIP headers. On outbound calls, the number is pulled from the active dialer queue record.
A: The system natively streams 16-bit linear PCM audio sampled at 16kHz or 8kHz (standard telephony narrow-band).
A: Yes. The dialer supports concurrent queue execution, subject to available SIP channels and configured API concurrency limits.
| Term | Definition in Sky Pro AI Calling Agent Context |
|---|---|
| Voice Activity Detection (VAD) | Algorithms that analyze audio streams to distinguish human speech from silence and background noise. |
| Streaming Speech-to-Text | Real-time transcription of audio streams into text as words are spoken. |
| Neural Text-to-Speech (TTS) | Deep-learning speech synthesis converting generated text into natural audio waveforms. |
| Barge-In | The ability for a human caller to interrupt the AI voice agent, instantly stopping synthetic speech. |
| Call Disposition | The final status assigned to a completed call (e.g., Hot Lead, Warm Lead, Demo Booked). |
| Turn-Taking Latency | The elapsed time from when a caller finishes speaking to when the agent begins its audio response. |
| SIP (Session Initiation Protocol) | Standard signaling protocol used for initiating, maintaining, and terminating VoIP telephone calls. |
| WebRTC | Open-source framework enabling real-time audio communication directly within web browsers. |
| SheetJS Parser | Client-side library parsing Excel spreadsheets into dialer queue records without server uploads. |