AWS networking interviews live in the details. Anyone can define a VPC. Fewer people can explain why a route table entry is missing, why a NAT gateway in the wrong availability zone doubles a data transfer bill, or when a Transit Gateway replaces a mess of peering connections.
Interviewers use this territory to separate people who've run production AWS networks from people who've read the documentation. Expect questions on VPC design, security groups versus NACLs, hybrid connectivity through Direct Connect, and Route 53 routing policies, often followed by a scenario where something isn't working and you have to say how you'd find out why.
These 20 questions cover that range, with sample answers specific enough to show real design and troubleshooting judgment instead of textbook definitions.
AWS networking at a glance
| Item | Details |
|---|---|
| Typical employers | Cloud consultancies, managed service providers, and any company running production workloads on AWS |
| Median pay | $134,050 a year for computer network architects overall (BLS, May 2025); BLS doesn't track AWS-specific networking roles separately |
| Job outlook | 8% growth from 2025 to 2035 for computer network architects, about 9,600 openings a year (BLS) |
| Education | A bachelor's degree in a computer-related field is typical, often with 5 or more years of networking experience |
| Certification | AWS Certified Advanced Networking Specialty, though AWS is retiring this exam with a last test date of December 31, 2026; the SAP-C02 Solutions Architect Professional exam covers overlapping VPC and hybrid connectivity material |
| Key tools | VPC console, Terraform or CloudFormation for provisioning, VPC Flow Logs, CloudWatch, Route 53, Direct Connect |
| Interview format | A technical screen on core VPC concepts, often a verbal or whiteboard design round, sometimes a scenario-based troubleshooting exercise |
How the interview usually works
- Recruiter or resume screen, confirming your AWS experience and whether you hold or are pursuing a networking-focused certification.
- Technical screen, typically 45 to 60 minutes, covering core VPC concepts and often a live troubleshooting scenario.
- Design round, where you're asked to design a VPC or hybrid network for a given set of requirements, sometimes on a whiteboard or shared document.
- Behavioral round with the hiring manager, often combined with the technical rounds at smaller companies.
General and background questions
1. What experience do you have designing or managing AWS networks in production?
Why they ask: They want to know the real scope of what you've run, not just which services you can name.
How to answer: Give a specific environment size, your tools, and what you actually own day to day.
Sample answer: I've managed the network layer for a multi-account AWS setup with about 40 VPCs across 3 regions, provisioned through Terraform and AWS Organizations. Day to day that means CIDR planning so accounts don't collide when they peer or attach to a shared Transit Gateway, writing security group and NACL rules for a PCI-scoped workload, and troubleshooting routing when a new subnet can't reach a shared services VPC. I also own our Direct Connect setup: 2 dedicated connections into different Direct Connect locations for redundancy, both terminating on a Direct Connect gateway that reaches VPCs in all 3 regions.
2. Walk me through how you'd design a VPC for a new production workload.
Why they ask: They want to see whether you plan CIDR ranges and subnet tiers before touching the console, or figure it out as you go.
How to answer: Cover CIDR sizing, subnet tiers across availability zones, and routing.
Sample answer: I'd start with a /16 for the VPC so we don't run out of address space later, then carve out /24 subnets: 2 public and 2 private across 2 availability zones for redundancy, plus a separate /24 for a database tier with no route to the internet. Public subnets get a route to an internet gateway; private subnets route outbound traffic through a NAT gateway per availability zone, since a single shared NAT gateway takes down the whole tier's outbound access if that zone has a problem. I'd add gateway endpoints for S3 and DynamoDB so that traffic never leaves the AWS network, and tag every subnet by tier and zone from the start.
3. Do you hold an AWS certification, or are you working toward the AWS Certified Advanced Networking Specialty?
Why they ask: Many networking roles list this credential as preferred, and they want your exact status.
How to answer: State where you actually stand, and show you know current details about the certification, not just its name.
Sample answer: I hold the Solutions Architect Associate and have been studying for the Advanced Networking Specialty, but I know AWS is retiring that exam, with the last test date on December 31, 2026, so I'm deciding whether to sit for it in the next few months or put that study time into the SAP-C02 Professional exam instead, which covers a lot of the same VPC and hybrid connectivity material. Either way, I keep hands-on with the material by rebuilding networking labs in a sandbox account rather than only working through practice questions.
Technical questions
4. What's the difference between a security group and a network ACL, and when do you use each?
Why they ask: This is a fundamental distinction, and a wrong answer here signals gaps in the basics.
How to answer: Cover stateful versus stateless, instance versus subnet level, and a reason to use both together.
Sample answer: A security group is stateful and attaches to an instance's network interface, so if I allow inbound HTTPS, the response traffic is allowed automatically. A network ACL is stateless and applies to the whole subnet, so I have to write matching inbound and outbound rules, and it evaluates rules in order by rule number until one matches. I use security groups for the everyday allow list, since they're easier to reason about, and I add a NACL rule when I need an explicit deny, like blocking a specific IP range flagged in a security incident, without touching every security group in the subnet.
5. How do you plan CIDR blocks across accounts to avoid problems later?
Why they ask: Overlapping CIDR ranges are one of the most common and hardest-to-fix mistakes in AWS networking.
How to answer: Describe a concrete planning habit, ideally with an example of what happens when it's skipped.
Sample answer: I check the CIDR range of every VPC I might ever peer or attach to a Transit Gateway before I pick one, since overlapping ranges block that connection outright and are painful to fix once resources are deployed. I default to non-overlapping /16 blocks per account, tracked in a shared spreadsheet so nobody reuses a range by accident. On one project we inherited a legacy VPC using 10.0.0.0/16, the same range as 3 other accounts, so we couldn't attach it to the shared Transit Gateway until we migrated its workloads into a new VPC with a unique range, which took about 6 weeks.
6. How does a Transit Gateway change network design compared to VPC peering?
Why they ask: They want to know you understand peering's limits, not just that Transit Gateway exists.
How to answer: Explain that peering doesn't transit, and describe how a Transit Gateway centralizes routing.
Sample answer: VPC peering is a 1-to-1 connection, and it doesn't transit, so if VPC A peers with B and B peers with C, A still can't reach C without its own peering connection to it. With 10 or more VPCs, that's a lot of individual connections to keep straight. A Transit Gateway acts as a single hub: every VPC attaches to it once, and I control routing between attachments with Transit Gateway route tables, including keeping a sensitive VPC isolated by putting it in its own route table. We moved from 12 peering connections to one Transit Gateway with 6 attachments, cutting the routing entries we had to keep in sync from around 40 to 6.
7. When would you choose AWS Direct Connect over a Site-to-Site VPN?
Why they ask: They want to see you weigh cost, setup time, and bandwidth needs rather than default to whichever sounds more advanced.
How to answer: Name the tradeoffs and describe how you'd design for redundancy if you did choose Direct Connect.
Sample answer: I'd use Direct Connect when a workload needs consistent bandwidth and latency that a VPN over the public internet can't guarantee, like continuous database replication to an on-premises data center. For a smaller or temporary connection, a Site-to-Site VPN is faster to stand up and cheaper. When I've set up Direct Connect, I've used 2 connections through different Direct Connect locations for redundancy, both terminating on a Direct Connect gateway so a single private virtual interface can reach VPCs in multiple regions, and I still keep a VPN as a backup path in case both Direct Connect connections have a problem at the same time.
8. What is AWS PrivateLink, and when would you use a gateway endpoint instead of an interface endpoint?
Why they ask: This tests whether you know the actual limits of each endpoint type, not just that VPC endpoints exist.
How to answer: Name what each endpoint type supports and a reason to pick one over the other.
Sample answer: PrivateLink lets an instance in a private subnet reach an AWS service or another VPC's service without going through an internet gateway or NAT device. Gateway endpoints only work for S3 and DynamoDB: they add a route in the subnet's route table pointing to the service, and there's no hourly charge. Interface endpoints work for most other AWS services, use a network interface with a private IP in your subnet, and charge per hour and per GB processed. I default to gateway endpoints for S3 and DynamoDB traffic since they're free and simple, and I use interface endpoints for anything else that needs to stay off the public internet, like Secrets Manager or an internal API another team shares through PrivateLink.
9. How would you troubleshoot an EC2 instance that can't reach the internet?
Why they ask: This is a common real problem, and they want an ordered process, not a guess.
How to answer: Walk through route tables, security groups, NACLs, and NAT or internet gateway health, in a logical order.
Sample answer: I check the subnet's route table first, since a missing or wrong route to an internet gateway or NAT gateway is the most common cause. Then I check the security group for an outbound rule allowing the traffic, and the NACL for both inbound and outbound rules, since NACLs are stateless and a missing outbound rule won't let return traffic back in even if the request left fine. If routing and rules look right, I check whether the instance has a public IP if it's in a public subnet, or whether the NAT gateway it depends on is healthy and in the right availability zone. I also pull VPC Flow Logs for the network interface to see whether traffic is being rejected at the NACL, the security group, or never leaving the instance at all, since the accept or reject field narrows it down fast.
10. What Route 53 routing policy would you use to fail over to a backup region during an outage?
Why they ask: They want to know you can match a routing policy to a real requirement, not just list the policy names.
How to answer: Name the failover policy and describe the health check that drives it.
Sample answer: Failover routing policy, paired with a Route 53 health check on the primary endpoint. I'd point the health check at a path that reflects real application health, like one that checks a database connection, rather than a plain ping, and set the secondary record with no health check so it's always available as the fallback target. I've used this for an API running primarily in us-east-1 with a secondary in us-west-2 behind a separate load balancer, with the health check interval at 30 seconds and 3 failed checks before failover, which kept failover fast without flapping on a single slow response.
11. How do VPC Flow Logs help you diagnose a network problem?
Why they ask: Flow Logs are the main tool for seeing what actually happened on the network, and they want to know you use them, not just recognize the name.
How to answer: Describe what Flow Logs capture and where you send them for different needs.
Sample answer: Flow Logs record accepted and rejected traffic at the network interface, subnet, or VPC level, including source and destination IP, port, protocol, and whether the traffic was accepted or rejected. I send them to CloudWatch Logs when I need to search quickly with Logs Insights during an incident, and to an S3 bucket when I want to keep them longer for a compliance review or query them with Athena. On one incident, Flow Logs showed a reject on a specific security group for traffic from an internal IP nobody recognized, which turned out to be a misconfigured health check coming from a load balancer in a different subnet than we expected.
12. What's the difference between a NAT gateway and an internet gateway?
Why they ask: People sometimes use these terms loosely, and the distinction matters for both cost and security design.
How to answer: State what each one is for and why you wouldn't use them interchangeably.
Sample answer: An internet gateway lets resources with a public IP send and receive traffic directly to and from the internet, so it's what a public subnet uses. A NAT gateway lets instances in a private subnet, which have no public IP, start outbound connections to the internet, like pulling an OS update, while blocking anything from starting a connection inbound. I put a NAT gateway in each availability zone rather than sharing one across zones, since routing private subnet traffic to a NAT gateway in a different zone adds a cross-zone data transfer charge and a single point of failure for that zone's outbound traffic.
13. How would you protect a public-facing application from application-layer attacks in AWS?
Why they ask: Networking roles often touch security too, and they want to know your approach beyond just security groups.
How to answer: Name AWS WAF and Shield, and describe how you'd tune rules instead of relying on defaults.
Sample answer: I'd put AWS WAF in front of the load balancer or CloudFront distribution, with managed rule groups for common threats like SQL injection and a rate-based rule capping requests per IP, tuned from real traffic patterns rather than a default number that blocks legitimate users. For volumetric attacks I'd rely on AWS Shield Standard, which covers every account by default, and consider Shield Advanced if the application is a likely target, since it adds cost protection during an attack and access to AWS's response team. I'd also make sure the origin, whether that's an application load balancer or an EC2 instance, isn't reachable directly, so an attacker can't get around WAF by hitting the origin's IP address.
Behavioral questions
14. Tell me about a network outage you diagnosed under pressure.
Why they ask: They want a real example of your troubleshooting process when something's actually broken and people are waiting.
How to answer: Describe the symptom, what you checked first and why, and the fix.
Sample answer: A production API started timing out for about 15% of requests during a deploy. I checked target group health first and saw instances cycling between healthy and unhealthy every couple of minutes, which pointed at the health check itself rather than the application. The health check path had changed in the new deploy, but the target group's health check configuration hadn't been updated to match, so it was hitting a 404 depending on which code version was live on a given instance mid-rollout. I fixed the health check path and the flapping stopped within 2 minutes. Afterward I added a deploy step that checks the health check config against the app's actual routes before a rollout ships.
15. Tell me about a time you had to explain a networking issue to a non-technical stakeholder.
Why they ask: Networking issues get escalated fast, and they want to know you can communicate clearly under pressure without talking down to someone.
How to answer: Describe what you told them, what you didn't know yet, and how that landed.
Sample answer: During a client-facing outage, our VP wanted an update every 10 minutes, but I only had partial information at first. I told him plainly what we knew, that a route table change had cut off a database subnet from the app tier, what we didn't know yet, whether it was a manual change or an automation error, and when to expect the next update, rather than guessing to fill the silence. That let him tell the client something accurate instead of a reassurance that might need correcting later. The fix took 40 minutes total, and afterward he said being told what we didn't know yet was more useful to him than a vague update would have been.
16. Tell me about a time you disagreed with a proposed network design.
Why they ask: They want to see you push back on a real risk with a specific reason, not just have an opinion.
How to answer: State the risk you saw, the alternative you proposed, and the outcome.
Sample answer: A colleague proposed a flat VPC design with every resource in public subnets to avoid NAT gateway costs. I pushed back because it left a database directly reachable from the internet if a single security group rule was ever misconfigured, which is a bigger risk than the NAT gateway cost, roughly $32 a month per gateway. I proposed private subnets for the database and app tier instead, with NAT gateways only where outbound internet access was actually needed, adding about $65 a month across the account. We went with the private subnet design, and it also let us pass a client security review that specifically asked whether any database had a public IP.
Situational questions
17. The number of point-to-point VPC peering connections between growing teams is getting hard to manage. What do you do?
Why they ask: This tests whether you'd recognize a peering setup outgrowing itself and know the standard fix.
How to answer: Describe migrating to a Transit Gateway and how you'd sequence the change safely.
Sample answer: I'd propose moving from full-mesh VPC peering to a Transit Gateway, since peering doesn't transit and the number of connections grows fast as more VPCs join. I'd map the current peering connections and CIDR ranges first to check for overlaps that would block the Transit Gateway attachments, then migrate VPCs a few at a time, keeping the old peering connections active until each migrated VPC's routing is confirmed working. I'd also set up separate Transit Gateway route tables so a team's VPC can't reach every other VPC by default, since a hub design without route segmentation just moves the same overly open access to a single choke point.
18. A partner needs access to one internal API, but security won't approve a VPN or public exposure. What's your approach?
Why they ask: This tests whether you know PrivateLink as a specific solution to a specific access problem.
How to answer: Describe setting up a VPC endpoint service and controlling who can connect to it.
Sample answer: I'd set up an interface VPC endpoint backed by a network load balancer in front of the API, and create a VPC endpoint service that the partner's AWS account can connect to with its own interface endpoint. That way the partner reaches the API over the AWS private network, with an endpoint policy limiting which principals can connect, and I'm not opening a security group to their IP range or standing up a VPN for a single API. I'd also set connection acceptance to manual on the endpoint service, so no outside account can attach without us approving it first.
19. A NAT gateway data processing bill has grown steadily, and you're asked to bring it down without cutting off internet access instances actually need. What do you do?
Why they ask: This tests whether you can find the real driver of a cost problem instead of just cutting service.
How to answer: Describe checking Flow Logs and Cost Explorer, then using endpoints to remove unnecessary NAT gateway traffic.
Sample answer: I'd start by checking Cost Explorer and VPC Flow Logs to see what's actually generating the traffic, since it's often one service pulling large files repeatedly rather than general instance traffic. If instances are calling S3 or DynamoDB through the NAT gateway, I'd add gateway endpoints for those, which are free and remove that traffic from the NAT gateway entirely. For traffic to other AWS services, I'd add interface endpoints where the monthly data volume costs less than continuing to route it through the NAT gateway. On one account, adding an S3 gateway endpoint cut NAT data processing charges by about 40% because a nightly backup job had been pushing several GB to S3 through the NAT gateway.
20. A new subnet can't reach a shared services VPC that's attached to your Transit Gateway. Walk me through how you'd find the problem.
Why they ask: This tests whether you have an ordered way to isolate a routing problem across multiple layers.
How to answer: Name the specific places you'd check, in order, and why.
Sample answer: I'd check 3 places in order: the subnet's route table, to confirm there's a route to the Transit Gateway for the shared services VPC's CIDR; the Transit Gateway route table associated with that attachment, to confirm it has a route back and that the new VPC's attachment isn't sitting in an isolated route table by mistake; and the security groups and NACLs on both ends, since a correct route doesn't help if a security group still blocks the traffic. On a similar case, the new VPC's attachment had been added to the Transit Gateway but never associated with the shared services route table, so traffic had nowhere to go past the Transit Gateway itself, which is easy to miss since the attachment shows as active either way.
Questions to ask the interviewer
- How many AWS accounts and VPCs does the team manage, and is routing built around a hub-and-spoke Transit Gateway or something else?
- What's the approach to CIDR planning and IP address management across accounts?
- Do you manage network infrastructure with Terraform, CloudFormation, or something else?
- How is Direct Connect or VPN connectivity structured for any hybrid or on-premises workloads?
- What does on-call look like for network issues, and how often does someone actually get paged?
- Is there budget or time set aside for AWS certification, including the Advanced Networking Specialty before it retires?
How to prepare
- Rebuild a multi-tier VPC by hand in a sandbox account: public and private subnets, a NAT gateway, route tables, and a gateway endpoint for S3, so you can talk through each piece instead of just naming it.
- Know the security group versus NACL distinction cold, including the S3 and DynamoDB-only limit on gateway endpoints.
- If you plan to bring up the Advanced Networking Specialty exam, know its status: AWS is retiring it, with a last test date of December 31, 2026.
- Review the Direct Connect Resiliency Toolkit and Transit Gateway documentation if the role touches hybrid connectivity.
- Practice explaining a routing problem out loud in the order you'd actually check it: route table, security group, NACL, then DNS.
If you're preparing for adjacent cloud roles, our cloud architect interview questions and solution architect interview questions guides cover broader AWS design topics, security engineer interview questions digs deeper into the security side of cloud infrastructure, and identity and access management interview questions is useful if the role touches IAM alongside networking.
Sources
- Amazon Web Services: aws.amazon.com/certification/certified-advanced-networking-specialty
- AWS Documentation: docs.aws.amazon.com/vpc/latest/userguide/how-it-works.html
- AWS Documentation: docs.aws.amazon.com/vpc/latest/privatelink/what-is-privatelink.html
- AWS Documentation: docs.aws.amazon.com/whitepapers/latest/aws-vpc-connectivity-options/aws-dire…
- AWS Documentation: docs.aws.amazon.com/Route53/latest/DeveloperGuide/routing-policy.html
- U.S. Bureau of Labor Statistics: bls.gov/ooh/computer-and-information-technology/computer-network…
by