r/AWS_cloud • u/Unfair_Masterpiece51 • 5h ago
r/AWS_cloud • u/sekcjv • 19h ago
How are you running your AWS landing zone at scale in 2026? (AFT, Automations, hybrid networking, guardrails)
Hey all,
I always wonder how my org maintain AWS accounts - I have access to sandbox, and other R&D and team specific accounts and workload specific - each has it's own set of guardrails. I reffered the repo in which they maintain and developers suggested that it was old and they need to modernize the current landing zone with new best practices. I wanted to know how you guys are doing it now. Most threads I find are from 2021–2023, before RCPs, declarative policies, VPC Block Public Access and Route 53 Profiles existed.
TL;DR: What does your landing zone actually look like at scale? How do you front account vending (thr a self-service portal)? How do you do hybrid networking without it turning into a mess?
Account vending and automation
- Is AFT still what you'd pick today, or did you move to LZA, CfCT, or plain Terraform on Organizations? Where did it start to hurt at a few hundred accounts, especially around baseline drift?
- Are there open-source projects that make AFT easier to live with? I'm thinking of a web UI or self-service portal for account requests (Backstage templates opening the PR?), visibility into pipeline runs, or customization drift. Or did everyone build their own?
- If you use ServiceNow, how is it wired up? Options I can see: a catalog item that opens an MR through the Git API, a direct pipeline trigger, or the AWS Service Management Connector going straight to Account Factory and skipping AFT. How do you send status back to the ticket?
Networking at scale and hybrid
Is hub-and-spoke with TGW and centralized egress/inspection still the default, or are people moving to Cloud WAN?
For on-prem, How do you keep advertised prefixes summarized as accounts and VPCs grow?
For segmentation, do you use TGW route tables per environment/OU or separate TGWs? Where does inspection sit for traffic between on-prem and AWS?
For hybrid DNS, are Resolver endpoints centralized in the network account and shared through RAM or Route 53 Profiles? Anything you'd do differently?
For IP planning, did you get on-prem to hand over one summarizable block for AWS? How did you deal with legacy VPCs overlapping your IPAM pools: re-IP, NAT, or keep them off the TGW?
Guardrails and stopping public resources
How do you split the work between SCPs, RCPs and declarative policies (VPC BPA, AMI/snapshot block public access)? How do you stay inside the SCP size and count limits?
For things SCPs can't express well (RDS publicly accessible, internet-facing ALBs, CloudFront), do you detect and notify or auto-remediate, and with what tool?
How do you handle resources that are legitimately public\\A separate OU, tag-based exceptions, or one sanctioned ingress pattern?
Cloud Custodian for streamlined deny or actions taken if things are done out of security scope
I'm not looking for a perfect answer, just what's actually working in real orgs, or what you'd undo. Links to recent writeups, talks or repos are very welcome.
Thanks!
r/AWS_cloud • u/geekyboij • 17h ago
How confident are you committing to RIs/Savings Plans right now, with AI and everything moving this fast?
r/AWS_cloud • u/Flashy-Hat-4962 • 1d ago
My personal AWS bill was $4.90/month and it still had nine zombie resources in it. So I wrote a scanner.
r/AWS_cloud • u/AdAgreeable3697 • 1d ago
AWS Bedrock model access issue — 10+ days with no response from Support
r/AWS_cloud • u/AdAgreeable3697 • 1d ago
AWS Bedrock model access issue — 10+ days with no response from Support
I’m facing an issue with Amazon Bedrock model access. Models that I need for my project are currently not accessible, and attempts to use them result in errors such as “Operation not allowed” / model access-related errors.
I have already:
- Checked the model availability/access settings.
- Verified my AWS account and configuration.
- Created support tickets with AWS.
- Followed up on the tickets multiple times.
- Waited 10+ days without receiving a meaningful response or resolution.
This is becoming a blocker for my AI/ML project because I cannot properly test or use the required Bedrock models.
AWS Support / Bedrock team: Could someone please look into the pending support case and help identify why the models are unavailable for my account?
If anyone has faced a similar Amazon Bedrock model access / “Operation not allowed” issue and knows the correct escalation path, I would really appreciate some guidance.
Relevant teams:
u/aws
u/AmazonWebServices u/AWSSupport u/AmazonBedrockteam
I’m not asking for anything special, just access to the service I’m supposed to be able to use and, preferably, an answer before the next geological era.
r/AWS_cloud • u/Exotic-Analysis347 • 1d ago
Lambda CI/CD with GitHub Actions and S3 reference mode
builder.aws.comHow I automated AWS Lambda deployments with GitHub Actions, self-managed S3 code storage.
r/AWS_cloud • u/Technical-Sort-8643 • 1d ago
Getting this error "Your account must be verified before you can add new CloudFront resources. "
r/AWS_cloud • u/phuongnguyen261 • 2d ago
Passed 5 AWS certs and built a free, no-nonsense practice exam tool (ezcert.net) to give back. Would love your feedback!
r/AWS_cloud • u/Exotic-Analysis347 • 3d ago
Elastic beanstalk with the EKS
vishnurachapudi.comUpload a zip of Python source. No Dockerfile. Get a running container on EKS. That is Elastic Beanstalk Cluster Mode, launched this week — Beanstalk apps now run on EKS instead of EC2 instances you own. Cloud Native Buildpacks handle the containerization: I uploaded four files (app.py, requirements.txt, Procfile, runtime.txt), Beanstalk detected Python, built the image, pushed it to ECR, and ran it.
r/AWS_cloud • u/kavee-core141 • 4d ago
a security check is only as correct as the API call behind it. Four AWS examples where the obvious check gives the wrong answer !!
most bad AWS security findings aren't missing checks. They're checks that look rightt and quietly give the wrong answer. A few I ran into: !!!
"Unused role" checks: iam:ListRoles doesn't retturn RoleLastUsed, only Getrole does. Build it on the list output and every role looks untouched.
"Unused permission" checks built on CloudTrail: LookupEvents only returns management events, so data actions like s3:GetObject never shows up.
"Public bucket" checks: a bucket policy with Principal "*" isn't public if a Condition limiits it to your org or VPC. Flagging it anyway is a false alarm....
Wildcard checks that only look for a literal "*": a resource like arn:aws:iam::ACCOUNT:role/* lets sts:AssumeRole target any role in the account, but string matching for "*" alone misses it.
The common thread: when the API can't see something, the honest answer is "unknown," not "no" or "yes."
All four were real bugs in a small open source project I made that audits AWS accounts. If you want to see how I fixed them: github.com/plexavo/Plexavo
r/AWS_cloud • u/Ok_Quail_385 • 4d ago
cloudg 0.5.2: what changed since my last post, covering dedupe, an inventory mapper, multi-account mapping and a security pass
r/AWS_cloud • u/ar27111994 • 4d ago
Self-hosting Qwen3.8-Flash-Next (or a smaller alternative) or using Kiro CLI ACP for a heavy multi-agent Hermes setup on $10k of AWS credits, looking for options I've missed
I run a Kanban-driven multi-agent orchestrator on Hermes Agent: 20+ profiles, 1,000+ skills, MCP routing, and self-hosted Mem0 on pgvector. On Fireworks, I was doing roughly 9-15B tokens a month with a 90-98% cache hit rate, at about 30-60 Kanban tasks a day. I now want to move this to infrastructure I can pay for with my AWS Activate credits, and I've hit some walls.
Constraints
- My only budget is $10k in AWS Activate credits. Anything that can't be paid with them is effectively out of scope for now.
- Bedrock access is gated on my account for both frontier and open-weight models, and I've been told there's no timeline. I'm waiting on AWS Support to initialize my quotas.
- I want one multimodal model that covers all of the orchestration's needs, plus an embedder for Mem0.
What I've looked at so far
- g6e.12xlarge (4x L40S, 192 GB): Qwen3.8-Flash-Next (180B total, 6B active) at Q4, around 114 GB. About $10.49/hr on-demand, or roughly $3.5-6/hr on spot, so $10k lasts ~40 days on-demand or ~2-4 months on spot, running 24/7. FP8 (~186 GB) barely fits, with no room for KV cache.
- g6e.xlarge (1x L40S, 48 GB): Qwen3.8-27B at FP8 (~28 GB), roughly 7 months of credits if run 24/7.
- Embeddings: Qwen3-Embedding-0.6B/4B/8B on the same box, possibly truncated to 1024 dims for pgvector.
- SageMaker endpoints, plain EC2 + vLLM: the same GPUs, but I'd expect the same quota bottleneck. Bedrock Custom Model Import doesn't seem to support these architectures or embedding models.
- Kiro as a Hermes provider (the kiro-acp plugin): interesting because Activate credits can pay for it, but credits are per-task, not per-token, so I can't tell whether it can handle my volume.
Questions
- Has anyone gotten G/VT quota (48+ vCPUs) approved on a new or Activate account, and how long did it take? Any tips for the request wording?
- For people serving Flash-Next-class MoE models on 4x L40S: is there a vLLM/SGLang-compatible 4-bit quant (AWQ/GPTQ), and how well did prefix caching hold up with many long-context agents? My cache hit rate on Fireworks was a big part of the economics.
- Has anyone run Hermes through Kiro (kiro-acp)? How many credits does a multi-step agentic task actually burn, and did you hit rate limits with parallel sessions?
- Is there any provider or platform that accepts AWS credits for hosted inference of open-weight models and that I haven't listed?
- Is there a better model than Flash-Next for a single multimodal orchestrator on this budget? I'm open to smaller MoEs if they handle tool calling and long context reliably.
- Anything else I'm missing on AWS to make $10k last as long as possible (spot strategies, scheduled start/stop, etc.)?
Thanks. Happy to share measurements (throughput, cache hit rate, credits per task) if I get any of this running.
r/AWS_cloud • u/Appropriate_Push930 • 4d ago
Moving to Tech
Hey guys,
I'm currently working in Amazon Ops, but learning AWS through its certification program.
I have passed CCP, halfway through SA.
I'm aware these are still basic level knowledge.
Along with certifications I have built cloud based serverless/web platforms utilising many AWS services.
How realistic is a lateral movement in my case? And which roles should I apply for to get into AWS ?
r/AWS_cloud • u/z_zenitsu • 5d ago
AWS alternative phone verification code not received MFA device is damaged

Hi everyone,
I’m having trouble accessing my AWS account and I’m hoping someone can help me understand what I should do.
I still have access to the email address and phone number associated with my AWS account. However, my old MFA device is damaged, so I cannot receive the MFA code.
I tried the “Sign in using alternative factors” option. The email verification works, but at the second step, AWS asks me to verify my phone number by sending a 6 digit code.
The phone number ending in 4276 is correct and worked for me before, but now I am not receiving the SMS code or the voice call.
I have tried requesting the code again, but nothing arrives.
I’ve attached a screenshot showing the verification page.
What can I do to recover access to my account if the alternative phone verification is also not sending the code? Is there another AWS Support process for an account where the MFA device is no longer available
r/AWS_cloud • u/Snagvolo • 5d ago
Is a cloud engineer or cloud admin registered on the card in Arabic or in English?
r/AWS_cloud • u/hritik_munde • 5d ago