Skip to content

Infrastructure

Building a Serverless CSV Search and Call Recording Platform with AWS

A serverless AWS platform I built to search and filter large CSV datasets stored in S3 and access customer call recordings, without using a database.

AWS Call Cover

Problem

The organization stored many CSV files in AWS S3. Those files were generated from their calls with customers. They also had a call recording for each record, which was stored in the same AWS S3 bucket.

Each CSV file contained a lot of information about the call and the customer. It included the customer name, agent name and ID, customer location, customer request, stored audio file name, customer email, and phone number. Each CSV file had hundreds to thousands of records.

The CSV files were organized in S3 using date-based prefixes. Inside each date-based folder were the CSV files for that date. The client needed a lot of filtering and searching options. They needed filtering by date range, and for a date range, there could have been more than a thousand CSV files, each containing a lot of data.

The client also needed to listen to and download those audio files and had plans for transcription as well. The audio files were linked to the CSV records using the file name stored in each record.

The application also needed user authentication through SSO with Amazon Cognito and role-based user access. Different users needed different access based on their user group. Some users should only be able to view the data stored in the CSV files. Some users should have the option to listen to the audio files only. Some users should have the option to both listen to and download those audio files.

AWS Call Problem

Constraints

The app had to be deployed on AWS. For the back end, we had to use Lambda functions, and all the files were stored in AWS S3.

The most important constraint was that no database could be used. We had to do everything using the existing CSV files. So, some of the filtering needed to happen at the S3 level, while other filtering and searching needed to happen inside the CSV files.

For most of the filtering, we also needed to read multiple CSV files into memory and combine their data before applying the filtering and searching.

Role-based authentication was also a must. For development, I had the option to work with mock data and a sandbox environment. However, for production debugging, I had to debug the issues together with the client because I did not have direct access to the production environment. The organization was large and had a lot of users, which made this more challenging.

AWS Call Constraint

Approach

For the front end, I chose Vite + React. The UI library was Tailwind CSS. I deployed the front end to an S3 bucket and used CloudFront to serve the application.

For the back end, I chose Python for the Lambda functions. The client was open to both Node.js and Python, but they favored Python because most of their existing application was written in Python. I also configured API Gateway to expose the backend APIs according to the project requirements.

As each request could involve working with a large number of CSV files, I was getting errors in the response. It was really hard to debug at that time because the error was not clear enough to immediately identify the issue.

I tried a lot of things to solve the problem. After a lot of debugging, research, and testing, I found that the issue was related to the Lambda resource allocation while processing the CSV files. I increased the memory allocation for the Lambda functions, and that resolved the error.

AWS Call Approach

Users were logging in to the application using SSO configured with Amazon Cognito. Cognito user groups were used to determine the access level for different users.

On the homepage, users were getting the data from the last seven days by default. There were a lot of filtering and search options available there.

For each request, the front end called the API through API Gateway. The request was authenticated, and the user's access was checked before processing the request. The frontend also checked the user's role to control which options were available to them, while the Lambda function checked the authorization again before processing the request.

The Lambda function first used the date-based S3 prefix to identify the relevant CSV files. It then read only those files and processed their data based on the requested filtering and searching options. For requests that required data from multiple CSV files, the files were read and combined in memory before applying the filters and search.

For the audio player, we looked for the corresponding audio file in the S3 bucket through the Lambda endpoint and returned it to the front end. The front end then played the audio using an audio player.

Outcomes

AWS Call Outcome

Previously, the organization had to maintain and work with all those files manually, so it was very hard to find a specific record. Finding the corresponding recorded audio was also a major hassle on top of that.

With the new application and its modern UI, it became much easier for them to manage those files and find records from the CSV data. It helped reduce their manual work and made it much easier for users to search through the call data and access the corresponding recordings.

Stack: Amazon Web Services (AWS) · AWS Cognito · Amazon S3 (AWS S3) · AWS Lambda · React · Tailwind CSS · Single Sign-on (SSO)

Contact

Let's work together

If you're hiring for a remote Staff or Principal role, or building something in healthcare, fintech, or AI, email me and we can talk it through.

Open to full-time and part-time remote roles

Also on Toptal. Email is the best way to start a conversation.