16 Cloud Security Tips and Best Practices to Protect Business Data
Cloud infrastructure often appears secure after deployment, then slowly drifts as teams ship features quickly and make rushed configuration changes. Most breaches come from small oversights that accumulate over time, such as exposed storage buckets, over-permissioned service accounts, or secrets pushed into version control.
Gartner projects that by 2025, nearly 99% of cloud security failures will stem from customer-side misconfigurations rather than provider failures. It reflects how fast cloud environments change and how many configuration options exist across any non-trivial deployment.
So in this guide, I share 16 cloud security tips drawn from what we at Agicent have learned while building and shipping cloud-hosted products across fintech, healthtech, and on-demand verticals since 2010.
I start with identity controls because every other layer depends on them, then move through infrastructure, data protection, monitoring, and finally the operational and human choices that decide whether those safeguards remain effective under pressure.
Identity: Starting Point Everything else Depends on
Cloud attackers typically don’t start by attacking infrastructure directly. They seek out valid credentials, since it takes fewer alerts and fewer traces to get authorised access than to get unauthorised access, and that is what they’re looking for. Identify controls that are effective under that sort of focused pressure are the base on which all other controls are built.
1. Apply Phishing-Resistant MFA to Privileged Accounts before Anything else
Multi factor authentication is understood correctly as a must. The issue that is often overlooked is the differences between MFA implementations. Codes sent via SMS and time-based one-time passwords are susceptible to being captured by current phishing kits that can pass off credentials as they are sent, during a targeted attack.
With FIDO2 and WebAuthn-based authentication, that moment of interception is overcome as the credential is cryptographically tied to the origin and the device, meaning it can’t be replayed even if the user is tricked into interacting with a phishing site.
The simple fact is it is applied to root accounts, accounts that have access to modify IAM policies and accounts that have access to production database before we even start any security work.
2. Scope IAM Roles to What the Task Requires
The reason IAM over-permissioning persists is not ignorance of the principle. A scoped IAM policy for a particular task will take longer to write than the attachment of a general managed policy; if a deployment is denied and the clock is ticking, a general policy will win.
The fix is not via a more stringent rule. It is a workflow change, which means that you will need to take the same effort as the broad one, but policy templates will be provided for common tasks that already have the minimum required permissions.
Just-In-Time access extends that; it issues elevated access permissions only for the time required for a specific task to be performed, and doesn’t maintain elevated access for anything longer than that.
3. Move Secrets out of Code and into a Managed Rotation Cycle
Documented and consistent: API keys and database credentials are exposed within minutes of being pushed to public repositories, and even private repositories don’t offer much more security than teams might think since repository access is typically more relaxed than access to the secrets stored inside.
A centralized secrets manager like AWS Secrets Manager or HashiCorp Vault eliminates the need to hardcode the secret into code entirely, and adds rotation to the credential life cycle. A rotated credential with a fixed expiration period reduces the impact of a compromise, if it occurs – since the window will close after a fixed time, without human intervention.

4. Regularly review and disable inactive identities
Clouds that have been operational for over a year have a growing pool of former employees, service accounts created for a project that was completed, and test users provisioned during a proof of concept. Active accounts are the ones that are most susceptible to monitoring, but the reason is that they are used frequently, thus making them the target of attacks by those that do not want the target account to be flagged by the monitoring system as a normal user.
So addressing the accumulation problem before it becomes a gap is done through a quarterly review that turns off any identity if there is no activity for the previous ninety days and an automated offboarding process that removes an identity’s access on the day that an employee is no longer employed or the project that they belong to ends.
Infrastructure: Configuration Decisions that Hold under Real Conditions
Cloud platforms are shipped with defaults geared to ease getting started, not production security. This is a fair product choice by the providers and a risk that engineering teams have to take into account. So now we focus on controls related to configuration choices where the difference between the default state and the secure state is greatest.
5. Know exactly where Your Responsibility Starts
The shared responsibility model outlines a cloud boundary of which the cloud provider and the customer are responsible for securing. Provider responsibility is for physical data centres, hardware, hypervisors and core networking fabric.
Customer responsibility includes operating systems, application configurations, access policies, workload security and data handling. It’s not at the boundary, of course, but in what teams think the provider should be able to deliver when they’re on the customer side.
Automated backups need to be configured and tested by the customer. With encryption at rest, customers must make a choice regarding key management. There is a customer process for patch management of the operating systems in customer managed instances.

6. Use CSPM Tooling to Catch Configuration Drift before Attackers do
Cloud Security Posture Management tools can do continuous scans of environments for adherence to security benchmarks and indicate when configurations are out of expected state.
The specific findings that are most worth focusing on are: Unencrypted storage volumes, exposed remote administration ports (e.g., RDP and SSH) without IP constraints, exposed storage buckets, and IAM roles that have wildcard permissions attached to public-facing resources.
As all of these categories are automatically produced by CSPM tools, this is important because the rate of change in cloud environments makes manual review in all but the smallest environments impractical.
While the risk prioritisation judgement (i.e. which findings to remediate first based on blast radius and/or exploitability) is still a human decision, the detection can and should be automated.
7. Secure Container Images before they Reach Production
Unvetted base images include the vulnerability profile of whatever is included in the base image at build time and can pass that on to the container images built on top of it. An image with no updates for a year and a half can have several dozen CVEs publicly known and exploits available to work.
The image scanning is integrated into a secure CI/CD pipeline implementation, allowing vulnerabilities to be identified before workloads reach production. The pipeline gate should prevent any images with critical severity findings from being forwarded to any environment above development, while allowing the image to be forwarded if a patch is not available for the finding.
8. Secure Container Images before they Reach Production
For the majority of cloud-based applications, APIs are the primary attack surface for connecting cloud services with mobile clients, third-party integrations and partner systems. Authentication, rate limiting, and input validation are the three controls that matter most and that are most consistently deferred until after an incident provides the motivation to implement them.
Rate limiting prevents abuse patterns, including credential stuffing and data scraping, from reaching the application layer. Input validation prevents injection attacks from reaching the database. Both are significantly easier to implement before an API has external consumers than after, because post-launch changes to API behaviour affect existing integrations.
9. Segment Networks and Limit Lateral Movement
A network where every workload can communicate with every other workload is a network where a single compromised resource provides a starting point for reaching everything else. Micro-segmentation divides the environment into smaller trust zones where communication between zones requires explicit rules rather than existing by default.
The practical implementation uses security groups, network ACLs, and private subnets to ensure that a web tier cannot directly query a database tier without passing through a defined network path, and that a compromised worker node cannot reach the secrets store without going through the same controls as any other consumer.
Cloud Security Controls: What Each Layer Protects Against
The table below maps security controls to the specific threat categories they address. Teams with limited capacity who need to prioritise should work top to bottom, since identity and configuration controls prevent the highest volume of incidents.
| Control Layer | Primary Threat Addressed | Failure Without It |
|---|---|---|
| Phishing-resistant MFA | Credential phishing and account takeover | Valid credentials provide full account access with no further barrier |
| Least privilege IAM | Privilege escalation and lateral movement | Compromised low-privilege account reaches high-value resources |
| Secrets management | Credential exposure via code or config | API keys and DB credentials accessible to anyone with repo access |
| CSPM scanning | Misconfiguration and configuration drift | Exposed storage, open ports, and overpermissioned roles go undetected |
| Container image scanning | Known CVEs in deployed workloads | Vulnerabilities with public exploits run in production undetected |
| Network segmentation | Lateral movement after initial compromise | Single compromised workload provides path to all other resources |
| API rate limiting | Credential stuffing, scraping, and abuse | Automated attacks run against endpoints at unlimited volume |
| Centralised logging | Undetected intrusion and slow incident response | No baseline for comparison and no signal when something changes |
Data Security: Protection that Starts with Knowing what You have
Encryption controls without data classification treat all data equally, typically resulting in the treatment of low-value data with the same controls that are used for regulated personal data, and gaps in alignment of classification and controls.
The sequence is important: Identify what you have, categorize by sensitivity, then set controls for each level.
10. Classify Data before Applying Controls
Cloud discovery tools can be automated to look through the cloud to determine where sensitive information resides in storage buckets, databases, and data pipelines. Classification provides structure by categorising data assets as public, internal, confidential or regulated.
After classification, policy enforcement can become more stringent, with access logging much more stringent, encryption key management more robust, and access controls much more limited, without having to review each asset manually.
The other approach, to impose uniform controls across the board, leaves the compliance gaps for data that is regulated and adds the overhead for data that really doesn’t need it.
11. Treat Encryption as a Key Management Problem
The ability to enable encryption on a storage volume or database is a checkbox that is part of the storage volume or database configurations on most cloud platforms. The key management decision determines the extent of protection that encryption really delivers. For many workloads, provider managed keys are convenient and appropriate.
Customer-managed keys in a Bring Your Own Key model ensure that the encryption key is under the control of the customer, even when the data is stored on the provider’s data infrastructure, for data that is regulated or where the encryption key is a meaningful competitive risk.
So a deliberate strategy is necessary to ensure TLS enforcement of data-in-transit – the security of data-in-transit depends on the intentional approach, just as it does with external data.
12. Address Shadow IT as a Data Flow Problem
Employees often share files or collaborate through unauthorised tools when approved platforms fail to meet operational needs quickly enough. In many organisations, this creates visibility gaps across the broader SaaS application ecosystem, allowing sensitive data to move outside governed environments without security teams being able to monitor it effectively.
Cloud Access Security Brokers emerge on the scene and identify unauthorized applications that are being used throughout the enterprise, expose the data that these applications are moving, and allow policy enforcement.
So the better solution is a process for evaluating, securing, and approving tools rapidly enough that the approved option becomes easier than the workaround.
This is especially important in modern SaaS application development environments where teams rely heavily on third-party integrations, shared workflows, and cloud-based collaboration systems. The process gives the incentive alignment which makes the visibility actionable, while CASB tooling provides the visibility itself.
Monitoring and Incident Response: Visibility that Produces Action
The difference between a contained incident and a full-scale breach lies in the detection capability. The teams that have the most controls do not necessarily have the most incidents that are reported quickly. It is they who can see what is normal and have rapid response channels if it isn’t.
13. Centralise Logs and Focus Attention on High-Signal Events
The logs in isolated services have no detection capability, since they must be correlated. The single view of activity through the centralisation of CloudTrail, Azure Monitor, application logs, and workload logs into a SIEM enables detection.
In this perspective, the events that should trigger immediate investigation are limited: for example, when the root account is used outside of a defined maintenance window, IAM privilege escalation that is not associated with an approved workflow, impossible travel events (two authentication events coming from geographically distant locations in a time frame that does not match the distance traveled, etc.), and bulk data access events (those that differ from the pattern of access events, etc.).
14. Use Agentless Scanning to Maintain Visibility without Friction
Security tooling with the demand of having to deploy an agent on every workload inflicts adoption friction which causes it to lose coverage over time. Teams deploy agents sporadically, don’t always put them into new workloads, and take them away when they start to cause performance issues.
Agentless vulnerability scanning software is based on analysing snapshots of virtual machines and container images, allowing the security team to gain visibility without altering the way engineering deploys workloads. The result is more uniform coverage since this tool does not provide an excuse not to work around it.
15. Automate First Response to Compress the Attacker’s Window
It is about closing the gap between the time an incident is detected and initial containment from minutes and hours to seconds. If the confirmed malicious IP’s access is attempted, an automatic firewall rule update may occur, blocking the access until the alert can be delivered to a human analyst.
If a credential becomes compromised, automated credential revocation can help narrow the access window until an investigation can ascertain how the credential was compromised. Regular tabletop exercises and game days will determine that these automated workflows behave as they are expected to behave in realistic situations, and if they don’t, they will fail in ways that only become apparent during real incidents.

Human Decisions that Determine whether Controls actually Hold
Technical controls are minimum standards. The baseline is determined by the decisions that people make on a day-to-day basis. A successful phishing attack on a cloud engineer with extensive production access, or an integration with a vendor that wasn’t fully specified, can tear up a well-built cloud environment. And they are the most frequent areas of attack.
16. Conduct Security Awareness Training Focused on Real Behaviour Change
Effective security awareness training shouldn’t be measured by the number of people who take an annual module.
It’s whether changes in behaviour occur over time in specific areas, such as handling of credentials, recognising and reporting on phishing attempts, and whether configuration changes are reviewed prior to implementation. Where click rates have been tracked over a number of months, as in a simulated phishing campaign, completion metrics do not give the same behavioural indication.
In fact, in development teams, training focused on secure coding practices in accordance to OWASP is beneficial since vulnerabilities that get to production are often patterns that are documented and avoidable with the correct practice.
How to Implement 3rd Party Vendor Risk Management as Your own Security Posture
Each vendor integration increases the attack surface by the extent of the vendor’s access. A vendor that has access to production data with a weaker security stance at their own organisation is an indirect exposure for a well-secured customer environment.
The practical method is to do two things: scope the vendor access that is just sufficient to enable the integration to work; understand the vendor’s incident response process prior to the integration going live.
It is common for access scope creep in vendor integrations because adding more access is easier than getting it right the first time and because of the pressure to get a vendor’s job done, this same effect happens.
Pre-Launch Cloud Security Checklist
The following table explains the individual verification tasks relevant to a cloud-hosted product, before it is first used by its users. These are decisions that should be made and should be explicitly confirmed, rather than assumed.
| Control | What to Verify Before Launch |
|---|---|
| MFA on privileged accounts | Root, IAM admin, and production DB accounts have phishing-resistant MFA enabled and tested. |
| IAM role scoping | No production role uses a wildcard resource or action where a scoped policy is feasible. |
| Secrets in vault | No credentials or API keys exist in environment files, code, or repository history. |
| CSPM baseline run | At least one full posture scan has run and critical findings have been remediated or documented. |
| Container images scanned | All production images have passed a CVE scan with no critical-severity unpatched findings. |
| Network segmentation confirmed | Database and secrets tiers are not directly reachable from public-facing workloads. |
| API rate limiting active | Rate limits are in place on all public-facing endpoints before external traffic arrives. |
| TLS enforced end-to-end | All external and internal service-to-service traffic uses TLS 1.2 or higher. |
| Centralised logging enabled | CloudTrail or equivalent is active and routing to a SIEM or centralised log store. |
| Automated alerts configured | Alerts for root usage, privilege escalation, and impossible travel are active and tested. |
| Inactive accounts reviewed | No accounts with no activity in 90 or more days retain active access. |
| Vendor access scoped | Each vendor integration has a documented access scope that has been reviewed and approved. |
How Agicent Approaches Cloud Security in Product Development
Architectural decisions come at a low cost, but once a product has been built and users are relying on it, they can be costly. From the initial sprint, Agicent’s engineering practice embeds security in the development process, with IAM design occurring at the architecture stage, secrets management as part of setting up a CI/CD pipeline, and CSPM tooling active before the first user interacts with the product.
The aim is not to provide a security review at the end of development but a set of engineering practices which are adopted throughout the build. It has been applied to a variety of products, including fintech, healthtech, and enterprise workflow automation products, and can be seen across our portfolio.
So, when designing cloud-based infrastructures, the Agicent MVP development practice is an obvious place to begin if a security approach is a part of the design. The most effective security controls to prevent the most incidents are the ones that go in, rather than the ones that are added after the first audit.
FAQs
How is cloud security different from securing on-premise infrastructure?
In two significant ways, the attack surface is unique. By design, cloud resources can be accessed from anywhere which does not translate to network perimeter controls. Cloud environments evolve more quickly than on premise environments, resources are being created and configurations are changed on an almost daily basis, and manual review is simply not a viable solution from a structural point of view. The most significant controls in a cloud environment are not network based, and are subject to rapid change, so they must be automated.
What does Zero Trust mean practically for a development team?
In the real world, Zero Trust translates to 3 engineering choices. The service-to-service communication is based on short-lived credentials and not long-lived keys. Unlike external APIs, internal APIs need to be authenticated by other internal entities. Network position is not the only thing that gives access; a workload residing within the private network must also provide valid credentials to access a database or secrets store. These are architectural decisions which can be made step-by-step, beginning with the highest-risk communication paths.
What are the most common cloud misconfigurations that lead to actual incidents?
Unauthorized access to open object storage buckets with public read access, overly broad IAM roles attached to public facing resources, open remote administration ports (no IP restrictions), and secrets in repository configuration files or environment variables make up a significant percentage of preventable incidents. All of these categories are automatically detected by CSPM tooling, so it's most effective to enable them before a product starts to be produced instead of after it becomes an issue.
How should a startup with limited security budget prioritise cloud security investment?
In order of impact: phishing-resistant MFA on privileged accounts, secrets manager integrated into CI/CD pipeline, at least one CSPM tool running in continuous scan mode and TLS enforcement on all traffic (including internal). These four controls are the most frequent categories of incidents, and can be put into place by an engineering team without a specific Security function. The rest can be added as the team and budget expand.
How to manage security in multiple cloud providers?
Multi-cloud security is a main issue of identity and logging. A central identity provider, using identity federation, guarantees consistent access rules across providers, as well as a change in one provider's access rules being propagated across all providers. Centralised logging is required for detecting across environments, as it brings together events from all providers in one view. It's more difficult to do across multiple providers, as each provider has its own set of tools, which is another reason to stick to two or three providers until the security operational maturity is reached that can handle multiple providers.