Short answer
RaftLabs built Vidmattic, a browser-based marketing video maker for a US client. Users stacked text, images, and gradients over a base video on a Konva canvas; the editor then serialised that layer stack into a timeline. Rendering ran serverlessly: the timeline went onto an AWS SQS queue, which triggered a Node Lambda with FFmpeg packaged as a Lambda layer. The Lambda compiled the timeline into an FFmpeg command, rendered the video, and wrote the output to S3, with the S3 path stored in Hasura and served from there. Every render had to complete inside the Lambda ceiling, which the team ran at its 900-second maximum. Active development ran from October 2020 to May 2021.
Engagement
The Vidmattic engagement
- Client
- Vidmattic, a US marketing technology product
- Sector
- Marketing technology and content creation tooling
- What it was
- A browser video maker for marketing videos. Start a canvas, drop in a base video, stack text, gradients, and images on top
- What we owned
- Four codebases: the Next.js editor, the NestJS and Hasura back end, a React admin panel, and the Lambda renderer
- Delivery
- Roughly six to seven months of active development, October 2020 to May 2021
- The engineering centre of gravity
- The render pipeline. Timeline JSON on SQS, FFmpeg packaged as a Lambda layer, output written to S3
- Evidence
- Engineering account only. RaftLabs holds no usage data, commercial record, or client-side history for this engagement
The situation
Vidmattic was an online video maker for marketing videos. You started a canvas, dropped in a base video, and kept stacking layers on top: text, gradients, images, and so on. The UI was deliberately plain. The whole point was that someone non-technical could put a marketing video together without opening an editing tool.
That description makes it sound like a front-end project. It wasn't. The editor is the part users see, and it's the part that took the least inventing. The hard problem sat behind the render button, where a stack of browser layers has to become a real video file without a render farm underneath it.
Our answer was to make the editor's output a data structure rather than a picture. The layer stack serialises into a timeline. That timeline goes onto a queue. A Lambda picks it up, compiles it into an FFmpeg command, renders the file, and writes it to S3. The database stores the path, and the video is served from there.
One thing to be straight about before you read further. The engineer who wrote the account this page is built on joined after the project had started, and he never held the client side of it. So this page covers engineering and nothing else. There are no outcome numbers here, no discovery story, and no budget shape, because we do not hold them. The evidence note near the bottom of this page says exactly what is missing and why.

Before and after
What the build had to make possible
- Making a marketing video meant opening real editing software, which is the specific thing the target user would not do
- Anything browser-based either previewed in the browser and stopped there, or needed servers sitting idle waiting for render jobs
- Video rendering is spiky by nature. Nobody renders steadily, they render in bursts, and a provisioned fleet is wrong in both directions
- A layer stack on a canvas and a finished MP4 are two different worlds with nothing obvious joining them
- A canvas editor where a base video takes text, gradient, and image layers, with no prior editing knowledge assumed
- The layer stack serialised into a timeline, so the editor's output is data the back end can reason about rather than a rendering of pixels
- A render path with no always-on compute: queue in, isolated Lambda per job, file out to S3
- FFmpeg available inside Lambda as a packaged layer, so the renderer itself stays small and deployable on its own
- A reusable translation layer between timeline JSON and FFmpeg arguments, so new layer types were an addition rather than a rewrite
- A local test setup that ran the whole pipeline on a developer machine
The problems worth reporting
Getting FFmpeg inside Lambda at all
FFmpeg is not a library you import. It's a binary, and a large one, and Lambda is not a machine you install things on. The options are to give up on serverless and run containers or EC2 with FFmpeg baked in, or to get the binary into the Lambda execution environment some other way.
We packaged FFmpeg as a Lambda layer. The renderer function then stays small and ships on its own, while the binary sits behind it as a separate, rarely changing artifact. That split is the reason deploys stayed quick on a project where the render logic changed constantly and the binary never did.
The cost is that layers are a real constraint, not a convenience. They have size limits, they are versioned separately from the function, and as the next challenge explains, they are also the first thing local tooling drops.
The 900-second ceiling is not a performance target, it's a wall
A Lambda invocation has a hard maximum duration. We ran the renderer at the 900-second maximum, which is the highest the platform allows. There's no tier above it. If a render doesn't finish inside fifteen minutes, it doesn't finish.
That changes how you think about the work. On a server you'd optimise a slow render because slow is annoying. Here, slow is a failure mode, so the FFmpeg command the Lambda builds has to be efficient by construction rather than tuned afterwards. Every layer type added to the editor is also a question about what it does to render time at the worst case.
If we were sizing this pipeline again, the honest engineering answer is that a hard ceiling with no escape hatch wants a fallback path for the jobs that would exceed it, whether that's a container-based renderer for long timelines or a segmented render that stitches parts together. On this project the ceiling held for the work in front of it.
Translating a timeline into an FFmpeg command without the code turning into string-building
FFmpeg is driven by command-line arguments, and complex compositions produce long, fragile argument strings. The naive version of this project has one function that takes the timeline JSON and concatenates its way to a command. That function grows with every layer type and becomes unreadable at about the third one.
We built a separate toolbox layer of reusable functions whose only job is turning pieces of timeline JSON into pieces of FFmpeg command. The renderer composes those pieces rather than building strings inline. That's the decision on this project worth copying: the mapping from editor concepts to encoder arguments deserves to be its own named thing, with its own deploy artifact, because it's where all the complexity accumulates.
Testing a layered Lambda locally, on tooling that doesn't support layers
Iterating on a render pipeline by deploying to AWS each time is slow enough to change how people work. So we ran LocalStack to get the pipeline on a developer machine.
The problem is that the free tier of LocalStack does not support Lambda layers, and layers were the whole basis of our packaging. Both FFmpeg and the toolbox lived in them.
We wrote a small converter script for local runs. It flattens the layers and the function into a single function, which LocalStack can execute. Developers got a working local pipeline, and the production packaging stayed as it was. This is a workaround and worth naming as one: local and deployed artifacts are not identical, so anything that depends specifically on layer resolution is not covered by local tests.
System flow
How a layer stack becomes a video file
The editor's job ends at producing a timeline. Everything after that is queue, compute, and storage, with no always-on server in the path.
Edit
Konva canvas, serialised to a timeline
A base video takes text, gradient, and image layers stacked on top, with state held in Recoil. On save, the editor turns that stack into a timeline: a data description of what sits where, over what, for how long. It is the contract between the editor and the renderer.
Queue
SQS
The timeline goes onto a queue rather than straight to compute. Submission is decoupled from execution, so bursts are absorbed instead of dropped.
Render
Node Lambda with FFmpeg as a layer
The Lambda builds the FFmpeg command from the timeline using the toolbox functions, then renders the video. It has 900 seconds, and no more.
Store
S3, with the path in Postgres
The output file is written to S3. The path is stored in the database behind Hasura, and the video is served from there.
the build
What we built
Four codebases, each with a clear job. The split matters because it's what let the render logic move quickly while the pieces around it stayed still.
A canvas editor that treats a video as a stack of layers, not a timeline of clips
The editor is built on Konva for the canvas and layer manipulation, with Recoil holding editor state and Grommet supplying the UI components. The mental model deliberately avoids the one professional editors use. There's no multi-track timeline to learn. There's a base video and things you put on top of it, which is a model a marketer already has from every other design tool they've used.

A GraphQL back end where the data layer is generated, not written
NestJS handles application logic over a GraphQL API, with Hasura sitting on Postgres as the data layer. Projects, layer stacks, render jobs, and output paths are all straightforward relational data, and generating the API over that schema meant the team spent its time on the render path rather than on writing resolvers for CRUD.

A renderer that is a data compiler, and is deployed as one
The Lambda's job is narrow: take a timeline, produce a command, run it, write the file. It lives in its own repository and deploys independently of the application. FFmpeg and the toolbox functions sit beneath it as layers, so a change to how a gradient is composited does not redeploy a video encoder.

A separate admin panel, kept out of the product
Operational tooling runs as its own React application rather than as privileged routes inside the user-facing product. That keeps admin capability out of the codebase users load, and it means the two can change on their own schedules.

What clients say
Most clients stay.
Some say so on camera.
Three-year average engagement. Founders and operators describing the work in their own words. No marketing varnish.
Olivia Martinez
Head of SMB Outreach, Vidmattic
The team's expertise in video software development transformed Vidmattic into an essential tool for our clients, streamlining video production and delivering exceptional results in engagement and ROI.
The lesson
In a media tool, the editor is the demo and the render pipeline is the product
If you're building browser-based media tooling, the thing most teams underestimate is which half of the work is hard. The editor gets the attention because it's what you show people. It's also the part with the most prior art: canvas libraries are mature, and layer manipulation in a browser is a solved shape. The render path is where the project actually lives or dies, and it's where the constraints are unforgiving in a way UI constraints are not.
The decision that made this pipeline workable was treating the editor's output as a data structure with a defined contract, not as something to be rendered directly. Once the layer stack is a timeline, the renderer is a compiler: timeline in, encoder command out. That framing is what let us put the translation logic in its own reusable toolbox and keep it out of the function that runs FFmpeg. Without it, the code that builds the command grows a branch for every layer type until nobody wants to touch it.
The constraint to design around from day one is the hard ceiling. Lambda gives you 900 seconds at most, and that number does not negotiate. Serverless rendering is the right call for spiky, bursty workloads where provisioning a fleet would be wrong most of the time, and it stops being the right call the moment a plausible job cannot finish in fifteen minutes. Decide in advance which side of that line your worst case sits on, and if it's close, build the fallback path before you need it rather than after a user finds it.
Where to go next
What we would look at next
These are observations about the architecture, not work that was scoped, commissioned, or delivered.
- Next 01
A fallback renderer for jobs that would exceed the ceiling
The 900-second limit has no escape hatch inside Lambda. A container-based renderer for long timelines, or a segmented render that stitches parts together, turns a hard failure into a slower success.
- Next 02
Close the gap between local and deployed artifacts
The LocalStack converter flattens layers into a single function so the free tier can run it. That's a sound workaround, and it means local tests do not exercise layer resolution. A thin deployed smoke test on each release would cover what local runs cannot.
- Next 03
Treat the timeline format as a versioned contract
The timeline is the interface between the editor and the renderer, and both sides evolve. Versioning it explicitly means an old saved project still renders after the editor has moved on.
- Next 04
Instrument the render path
Queue depth, render duration against the ceiling, and failure reasons are the three numbers that tell you whether this architecture is still the right one. On a pipeline with a hard limit, the distribution of render times matters more than the average.
How it runs
- AWS Lambda
- Render load is bursty, so a provisioned fleet is either idle or short. Lambda runs each job in isolation with nothing running between jobs. The trade we accepted for that is the 900-second maximum duration, which we ran at its ceiling and which has no tier above it. FFmpeg ships as a Lambda layer, which keeps the renderer function small and independently deployable.
- AWS SQS
- The queue sits between the editor and the compute so job submission is decoupled from job execution. That gives natural back-pressure on bursts, and makes retries and dead-letter handling a configuration concern rather than something the application has to implement.
- AWS S3
- Rendered files are written to S3 by the Lambda, and the object path is stored in Postgres. Video files are large and rarely modified after creation, which is exactly what object storage is for. No separate media server sits in the path.
- Next.js
- The editor is the application, so the front end needed to be a real app rather than a rendered page, while the surrounding marketing and account surfaces wanted fast first loads. Next.js covers both without a second front-end stack.
- Konva
- The product's core interaction is manipulating stacked layers on a canvas. Konva gives that model directly, with hit detection, transforms, and layer ordering handled, instead of building canvas primitives from scratch. Recoil holds the editor state around it, and Grommet supplied the surrounding UI components.
- NestJS
- The back end coordinates projects, render jobs, and outputs over GraphQL. NestJS gave the application layer a structure that stayed navigable as the render path grew more involved.
- Hasura
- Projects, layer stacks, job status, and output paths are ordinary relational data. Hasura generates the GraphQL API over the Postgres schema, so the team wrote resolvers for the things that needed logic and none for the things that did not.
- Firebase
- Authentication was not where this project's difficulty lived, and building it would have taken time from the render pipeline. Firebase handled sign-in with no infrastructure for the team to run.
Common questions
Render load is bursty. Nobody renders at a steady rate, so a provisioned fleet is either sitting idle or too small during a spike. Lambda runs each job in isolation with nothing running between jobs, and SQS in front of it absorbs bursts. The trade is a hard 900-second maximum per invocation with no tier above it, so serverless is the right answer only when your worst-case job comfortably fits inside that.
Package it as a Lambda layer. FFmpeg is a binary rather than an importable library, and Lambda is not a machine you install software on, so the binary ships as a separate versioned artifact that the function sits on top of. On Vidmattic this also kept deploys fast: the render logic changed constantly and the encoder binary never did, so only the small function needed redeploying.
By making the editor's output data rather than pixels. The Vidmattic editor serialised its layer stack into a timeline, which is a description of what sits where, over what, for how long. That timeline is the contract between editor and renderer. The renderer then behaves like a compiler: timeline in, FFmpeg command out, file to S3. Keeping the translation logic in its own reusable set of functions is what stops it turning into unmaintainable string-building.
Partly. We ran LocalStack to get the pipeline onto developer machines, but its free tier does not support Lambda layers, which was the whole basis of our packaging. We wrote a converter script that flattens the layers and the function into a single function for local runs. It works, and it means local and deployed artifacts are not identical, so anything depending specifically on layer resolution needs a deployed check.
Roughly six to seven months of active development, from October 2020 to May 2021. That figure comes from repository history rather than from a contract, and it covers four codebases: the editor, the back end, the admin panel, and the Lambda renderer.
We are not publishing either, because we do not hold them. The engineer whose account this page is built on joined after the project started and never held the commercial side, and no one currently at RaftLabs has that record. We would rather say that than put a range on the page. For what a comparable build costs today, talk to us and we will scope it properly.
Because we cannot stand behind any. A previous version of this page carried three metrics with no source behind them, and they have been removed. RaftLabs never held analytics access for this product, and the engagement ended in 2021. Every other claim on this page traces to a named source, which is the standard we would rather be held to.
If your product turns user-assembled state into a rendered media file and your load is spiky, the shape transfers directly: a defined serialisation format, a queue, isolated compute, object storage, and a database that only holds the path. The question to answer first is your worst-case render duration. If it fits inside fifteen minutes with room to spare, this architecture is a good fit. If it does not, you want a container-based renderer or a segmented approach instead.
Related work


