← Notes from the workbench

What is Amazon Connect?

A plain explanation of Amazon Connect: what it does, the handful of words you need to know, and the two places a custom AI voice agent can plug in. Background reading for the KVS build write-up.

· 6 min read

Amazon Connect is a phone system for businesses, run by AWS.

That really is the whole idea. You rent phone numbers from Amazon, you describe what should happen when somebody calls one of them, and Amazon gets the call to a person who can help. It handles chat and a few other channels too, but voice is the core of it, and voice is what this piece is about.

I am writing this as background. I have a longer write-up on building an AI voice agent on top of Connect using Kinesis Video Streams, and that piece assumes you already know what the moving parts are called. This is that assumed knowledge, in plain language.

One note on the name first, because it trips up almost everybody. It is Amazon Connect. You will hear people say AWS Connect or Amazon Cloud Connect, and they mean this. What they do not mean is AWS Direct Connect, which is a completely unrelated networking service that runs a private line into your VPC. Search for the wrong one and you will end up reading about fibre cross-connects, wondering where the phone calls went.

What happens when somebody calls

The clearest way to understand Connect is to follow one call from ring to hang-up.

  1. Someone dials your number. You claimed that number inside Connect, from Amazon’s stock, in whichever country you need. Amazon deals with the phone carriers. You never talk to one.
  2. Connect answers and starts running a contact flow. Every number has a flow attached. The flow is what happens next.
  3. The flow is a little program you draw in the browser. You drag blocks onto a canvas and join them with arrows. Blocks do things like play a message, collect keypresses, look something up, or decide where to send the call.
  4. Speaking to the caller is a block. You type the sentence and Amazon Polly reads it aloud, so you are not recording and re-recording audio files every time the wording changes.
  5. Looking something up is a block too. That block calls an AWS Lambda function. Your code runs, returns some JSON, and the flow branches on the answer. This is where “is this caller overdue on their bill” gets decided.
  6. If the flow cannot finish the job itself, it puts the caller in a queue. A queue is a waiting line, usually one per skill or department.
  7. A routing profile decides who gets it. Every agent has one. It says which queues that agent takes calls from, and which of those come first when several are waiting.
  8. An agent’s softphone rings. Not a desk phone. A browser tab, using WebRTC. They click accept and they are talking to the caller.
  9. The call ends and Connect writes it down. A recording goes to S3 if you enabled it, and a structured summary of everything that happened goes to Kinesis.

That is the product. Numbers at the front, a flow in the middle, agents at the back, and a record of it afterwards.

The words you need to know

Six terms cover most conversations about Connect.

  • Instance. Your own private copy of Connect. It has its own web address, its own users, its own flows and numbers. Most teams run a separate one per environment.
  • Contact. Any single interaction, whether that is a voice call, a chat, or a task. Each one gets a contact ID that follows it through every system it touches. When you are debugging, this is the string you search for.
  • Contact flow. The drag-and-drop program from step 3 above. Sometimes just called a flow.
  • Queue. The waiting line callers sit in before an agent picks up.
  • Routing profile. Attached to an agent. Decides which queues they serve and in what order.
  • CTR, or Contact Trace Record. The structured record written when a contact ends. Who called, what the flow did, how long they waited, which agent took it, what your Lambda functions stored along the way. Send these to your warehouse early. The first time somebody asks how many callers hung up while waiting, you will want a year of history rather than a week.

Where AI fits in

There are two doors into Connect for anything intelligent, and choosing between them decides most of your architecture.

The front door is Amazon Lex. It is built in. A flow block hands the conversation to a Lex bot, the bot listens to the caller, works out which intent they matched, collects any details it needs, and hands a result back so the flow can branch on it. For “tell me why you are calling” and for capturing structured things like an account number, it works and it is quick to set up.

The limits show up fast. Lex matches intents from a list you wrote in advance, which is not the same thing as holding a conversation. You do not choose the speech recognition, you do not get the raw audio, and you cannot put your own model in the middle. Each hop between the flow, the bot and your Lambda also costs time, and on a phone call people hear that time as hesitation.

The side door is Kinesis Video Streams, usually shortened to KVS. A “start media streaming” block tells Connect to publish the live audio of the call to a video stream while the call is still happening. Your own service reads from that stream and does whatever you want with it: your choice of speech recognition, your model, your logic, your voice.

What the KVS path actually gives you

Worth being precise, because this is where the build write-up picks up.

  • Two separate audio tracks. AUDIO_FROM_CUSTOMER is what the caller says. AUDIO_TO_CUSTOMER is what they hear, which includes the agent, the system prompts, and any Lex bot in the flow. They stay separate, which spares you the job of pulling two voices apart. You pick one direction or both when you configure the block.
  • Telephone-grade audio. 8 kHz, 16-bit signed PCM, mono, little-endian. That is a low sample rate by modern standards. Speech recognition copes fine, but check it against whatever you plan to feed it, because plenty of models expect 16 kHz.
  • A way to find the stream. When streaming starts, the flow writes contact attributes holding the stream ARN, the fragment number to start reading from, and start and stop timestamps. Read those attributes and you know exactly where in the stream your call begins.

Now the part that catches people out, and the reason building a voice agent this way is more work than it first appears. Media streaming is a one-way door. It gets audio out of the call. It does not give you a way to push audio back in.

You can listen to the caller in real time, transcribe them perfectly, and generate a brilliant reply, and still have no obvious route for that reply to reach their ear. Working out the return path is the real engineering problem here, and it is the thing worth writing about.

The short version

Connect is phone numbers, a visual flow that decides what happens on a call, queues and agents at the far end, and a trail of records afterwards. Lex is the easy way to add intelligence and it runs out of room quickly. KVS hands you the raw audio and hands you the hard problem along with it.

That hard problem is what the build write-up is about.