All talks

From Kafka to S3: Streaming Data Pipelines with Kafka Connect

AWS User Group Hyderabad45 min talk, live demo and Q&A

  • Kafka
  • Kafka Connect
  • S3 Sink
  • AWS
Ashfaq at a lectern beside a screen showing a resources slide, speaking to a room of seated attendees

How do you get streaming data from Kafka into Amazon S3 without writing and babysitting custom consumer code? Kafka basics, what Kafka Connect is and why it exists, then a deep dive into the S3 Sink connector from someone who works on it, ending with a pipeline built live.

What I covered

I opened with the problem every growing system hits, a tangle of point-to-point integrations, and showed how Kafka in the middle untangles it. I kept the Kafka basics short and spent most of the time on the S3 Sink connector, with real configs and the mistakes I see while working on it.

  • Kafka fundamentals: the core ideas, enough to follow the rest of the talk.
  • Kafka Connect: what it is, and why it beats writing your own integration code.
  • S3 Sink connector: landing Kafka data in Amazon S3, and the settings that matter.
  • End to end: a sample IoT pipeline, from devices to analytics on AWS.

Live demo

  1. A Kafka broker and a Connect worker running in Docker Compose.
  2. Create an S3 bucket, then create the S3 Sink connector with one REST call.
  3. Produce JSON orders and watch the objects appear in S3.

Gotchas I shared

  • flush.size: too small means many tiny S3 objects and high PUT costs; too large delays writes. Each file is one PUT request.
  • rotate.interval.ms vs rotate.schedule.interval.ms: record time vs wall-clock time. Know which one you need.
  • Use Parquet for analytics: columnar and compressed (snappy, gzip or zstd), and Athena reads it well.
  • Enable a dead-letter queue so bad records don’t block the pipeline.
  • IAM permissions: the connector needs PutObject, GetObject, AbortMultipartUpload, PutObjectTagging, ListAllMyBuckets, ListBucket and GetBucketLocation.
  • Schema compatibility: set schema.compatibility (BACKWARD, FORWARD or FULL) to catch breaking changes.

More talks

Ashfaq presenting a slide titled "Kafka: where every order lives", showing three brokers replicating partitions, to a seated audience

Apache Kafka Meetup Hyderabad

Kafka, Demystified: How Your Swiggy Order Reaches Everyone Who Needs It

One food order has to reach the restaurant, a delivery partner, payments and your notifications. This talk follows that single order through Kafka, Kafka Connect, Schema Registry, Kafka Streams and Confluent Cloud, so they read as one connected system instead of five tools, and builds the pipeline live.

  • Kafka
  • Kafka Connect
  • Schema Registry
  • Kafka Streams
  • CDC

Companion post