Getting Started with AWS CDK for Data Platform Teams — A Practical Guide
In this article let us look at AWS CDK — the Cloud Development Kit — and how data platform teams can use it to manage their infrastructure. If you have been using Terraform or CloudFormation to spin up your data stack, CDK gives you a way to do the same thing but using actual programming languages like TypeScript or Python.
When I first started building data pipelines on AWS, everything was manual. I would go into the console, create an S3 bucket, set up a Glue database, attach IAM policies — and then hope I remembered the steps when I had to do it again for a different environment. That approach does not scale. CDK solves this by letting you define your infrastructure in code, with all the loops, conditions, and abstractions that come with a real programming language.
This article walks through setting up CDK from scratch and provisioning a few resources that a typical data platform would need — an S3 bucket for raw data, a Glue database for the catalog, and the IAM roles to tie them together. By the end, you should have enough to start writing CDK for your own data pipelines.
Why CDK Over Terraform or Raw CloudFormation?
If you are already using Terraform, you might wonder why bother with CDK. Here is how I see it:
| Approach | Best For | Watch Out For |
|---|---|---|
| Raw CloudFormation | Simple setups, few resources | YAML gets verbose fast. No logic — lots of copy-paste |
| Terraform | Multi-cloud teams, large orgs with existing HCL code | HCL is its own language. Loops and conditionals feel awkward |
| AWS CDK | AWS-only teams who want code reuse, testing, and familiar languages | Tightly coupled to AWS. CloudFormation is still under the hood |
CDK synthesises CloudFormation under the hood, so you get the reliability of CloudFormation with the expressiveness of TypeScript or Python. For data teams that are already writing Python or TypeScript for their pipelines, CDK means you do not need to learn another DSL just to manage infrastructure.
1. Prerequisites
Before we start, make sure you have these installed:
- Node.js 18 or later (CDK runs on Node even if you write stacks in Python)
- AWS CLI configured with a profile (
aws configure) - An AWS account with enough permissions to create S3, Glue, and IAM resources
Install the CDK CLI globally:
1
npm install -g aws-cdk
Verify the installation:
1
cdk --version
2. Bootstrapping Your AWS Account
CDK needs a one-time bootstrap in each AWS account and region you plan to deploy to. This creates a CDK toolkit stack — basically an S3 bucket and some IAM roles that CDK uses during deployment.
1
cdk bootstrap aws://ACCOUNT-NUMBER/ap-southeast-2
Replace the account number and region with yours. You only run this once per account-region pair. If you skip this step, CDK will remind you with a clear error message when you try to deploy.
3. Creating a CDK Project
Create a new directory and initialise a CDK app. I will use TypeScript here, but the concepts carry over to Python.
1
2
mkdir data-platform-cdk && cd data-platform-cdk
cdk init app --language typescript
This gives you a folder structure with a bin/ directory (the entry point) and a lib/ directory (where your stacks go). The generated lib/data-platform-cdk-stack.ts is your starting point.
4. Writing Our First Data Platform Stack
Let us replace the boilerplate stack with something useful. Open lib/data-platform-cdk-stack.ts and write the following:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
import * as cdk from 'aws-cdk-lib';
import * as s3 from 'aws-cdk-lib/aws-s3';
import * as glue from 'aws-cdk-lib/aws-glue';
import * as iam from 'aws-cdk-lib/aws-iam';
import { Construct } from 'constructs';
export class DataPlatformStack extends cdk.Stack {
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
// S3 bucket for raw data landing zone
const rawBucket = new s3.Bucket(this, 'RawDataBucket', {
bucketName: `raw-data-${this.account}-${this.region}`,
versioned: true,
encryption: s3.BucketEncryption.S3_MANAGED,
lifecycleRules: [
{
id: 'TransitionToGlacier',
transitions: [
{
storageClass: s3.StorageClass.GLACIER,
transitionAfter: cdk.Duration.days(90),
},
],
},
],
});
// Glue database for the data catalog
const glueDb = new glue.CfnDatabase(this, 'RawDataCatalog', {
catalogId: this.account,
databaseInput: {
name: 'raw_data_catalog',
description: 'Catalog for raw ingested data',
},
});
// IAM role for Glue jobs to read from S3 and write to the catalog
const glueRole = new iam.Role(this, 'GlueJobRole', {
assumedBy: new iam.ServicePrincipal('glue.amazonaws.com'),
managedPolicies: [
iam.ManagedPolicy.fromAwsManagedPolicyName(
'service-role/AWSGlueServiceRole'
),
],
});
rawBucket.grantReadWrite(glueRole);
// CloudFormation outputs
new cdk.CfnOutput(this, 'RawBucketName', {
value: rawBucket.bucketName,
});
new cdk.CfnOutput(this, 'GlueDatabaseName', {
value: glueDb.ref,
});
}
}
Let me break down what is happening here.
The S3 bucket: We create a versioned bucket for the raw data landing zone. The lifecycle rule moves objects to Glacier after 90 days — which makes sense for a raw zone where old data is rarely accessed but cannot be deleted due to compliance. Notice we use this.account and this.region in the bucket name because S3 bucket names must be globally unique.
The Glue database: Glue databases in CDK use the L1 construct (CfnDatabase), which maps directly to the CloudFormation resource. There is no L2 construct for Glue databases at the time of writing. This means less syntactic sugar — you pass the properties exactly as CloudFormation expects them.
The IAM role: We create a role that Glue jobs can assume, attach the AWS-managed Glue service policy, and then grant it read-write access to our S3 bucket. The grantReadWrite method is CDK magic — it automatically generates the right IAM policy statements without you having to write ARN patterns by hand.
5. Synthesising and Deploying
Before deploying, run the synth command to see what CloudFormation template CDK generates:
1
cdk synth
This outputs the CloudFormation template to stdout. It is worth looking at it once — you will see how the TypeScript code translates into YAML. It also helps you catch mistakes before anything touches AWS.
If everything looks good, deploy:
1
cdk deploy
CDK will show you a diff of changes and ask for confirmation. Type y and let it run. After a minute or two, you should see the stack outputs printed in the terminal.
1
2
DataPlatformStack.RawBucketName = raw-data-123456789012-ap-southeast-2
DataPlatformStack.GlueDatabaseName = raw_data_catalog
Things to Be Careful About
Here are a few things I noticed while working with CDK in data projects.
L1 vs L2 constructs: Not every AWS service has high-level L2 constructs. Glue databases and tables, for example, still use L1. This means you need to understand the underlying CloudFormation properties. The AWS CDK API reference is your friend here.
State management: CDK uses CloudFormation under the hood, so you inherit CloudFormation’s state behavior. If you delete a resource from your CDK code, CloudFormation may try to delete the actual resource on the next deploy. Use removalPolicy: RemovalPolicy.RETAIN on S3 buckets and databases if you want to keep data when the stack is destroyed.
Bootstrapping across accounts: If your team manages multiple AWS accounts (dev, staging, prod), you need to bootstrap each one. CDK supports cross-account deployments, but the permissions model gets involved. Plan your account structure early.
CDK version upgrades: The CDK team ships updates frequently, and sometimes APIs change between versions. Pin your CDK version in package.json rather than floating on latest, especially in CI/CD pipelines.
What Changes in a Production Use Case
The stack above is a decent starting point, but for production you would want to add a few things:
- Encryption with KMS: Replace
S3_MANAGEDencryption with a customer-managed KMS key. This gives you control over key rotation and access policies. - Separate stacks per layer: Split your infrastructure into multiple stacks — one for storage (S3), one for catalog (Glue), one for compute (Glue jobs, EMR). This keeps blast radius small.
- CI/CD integration: Run
cdk deployfrom GitHub Actions or CodePipeline. The CDK has acdk pipelineconstruct that makes this straightforward, though that deserves its own article. - Testing: CDK has a
fine-grained assertionslibrary. Write tests to check that your stack has the right bucket policies and IAM roles before deploying. - Naming conventions: In production, define a naming convention and wrap it in a helper. Something like
${project}-${environment}-${resource}keeps things sane when you have dozens of buckets.
CDK gives data teams a way to manage AWS infrastructure without switching between tools and languages. If your team is already writing Python or TypeScript for pipelines, CDK fits naturally into the workflow. It is not a silver bullet — you still need to understand the AWS services you are provisioning — but it removes a lot of the boilerplate and makes your infrastructure testable and reusable.
In a future article, I will cover how to use CDK Pipelines to set up a full CI/CD flow for data infrastructure. Until then, the stack above should be enough to get you started.

