September 29, 2026 · 2 min read
Store the original first: an email pipeline on SES, S3 and Lambda
Processing incoming email straight in Lambda was simpler, but it tied the raw message to the processing step. Saving every .eml to S3 first makes processing disposable: change the code, fix a bug or rebuild derived data, then replay the originals.

The design this grows into: a parser that writes metadata to DynamoDB and attachments to S3, and a small app to read the mail. The raw store is what makes adding those pieces safe.
Kustto needed addresses on its own domain, like hola@ and proveedores@, without paying for a mailbox per address. Amazon SES can receive mail for a domain, run actions on each message and hand it to a Lambda function. The first version of the design was exactly that: SES invokes Lambda, Lambda reads the message and processes it.
It was simpler, and it had a problem I only noticed while writing the processing code. The original message existed only inside that one invocation. If the code had a bug, or I wanted to change how messages were handled later, there was no copy of the input to go back to.
One more action in the rule
SES receipt rules run their actions in order. The fix was to make the first action store the raw message in S3 and the second one invoke the function:
{
"Name": "reenviar-todo",
"Recipients": ["kustto.com.mx", ".kustto.com.mx"],
"Actions": [
{ "S3Action": { "BucketName": "kustto-inbound-<account>-us-east-1",
"ObjectKeyPrefix": "entrada/" } },
{ "LambdaAction": { "FunctionArn": "arn:aws:lambda:us-east-1:<account>:function:kustto-correo",
"InvocationType": "Event" } }
]
}Now every message lands in a private bucket as the complete .eml, exactly as it arrived, before any code touches it. The object key is the SES message ID, so the function can find it from the event alone.
What the function does today
Today the processing step forwards each message to a team inbox. The function reads the raw message from S3, rewrites the headers that would make the forward fail, and sends it with SES:
- It removes the original DKIM, ARC and authentication headers. They are signed for the sender's domain and would fail once the message is sent from Kustto's.
- The new
Fromis a verified Kustto address with the sender's name in it, andReply-Tokeeps the original sender, so replying just works. - It reads the recipient from the SMTP envelope instead of the
Toheader, because a message that arrived as a blind copy doesn't list the address it was sent to. That recipient goes in anX-Kustto-Paraheader, the only way to tell hola@ from ventas@ when the whole domain is accepted.
const destinatarios = registro?.ses?.receipt?.recipients ?? [];
const para = destinatarios.length > 0
? destinatarios.join(", ")
: valorDe(cabeceras, "to");If sending fails, the function throws. Because it is invoked asynchronously, Lambda retries it, and since the message is already safe in S3, a retry costs nothing and loses nothing.
Why the order matters
With the original stored first, processing becomes disposable. I can rewrite the function, fix a bug in how headers are handled or add a completely new step, and run it again over the messages already in the bucket. Nothing depends on what happened during the first execution.
That is also what makes the next step easy. A parser that extracts the sender, subject and attachments into DynamoDB, and a small app to search them, can be built and backfilled from the bucket. If its schema turns out wrong, the table can be dropped and rebuilt from the same originals.
It is a very small change in the architecture, one extra action in a rule and a bucket, but it turns every bug in the processing step from data loss into a re-run.