Data risks when using AI do not lie in the technology itself, but in the habits of the team and the absence of internal principles.
The most significant risk is personnel inputting sensitive data (contracts, client information, source code) into public AI tools, where data may be stored or used to train models. This is not a theoretical concern: there have been large companies that tightened internal AI usage rights after discovering personnel pasted source code into public tools. The solution is not to ban AI, but to manage it as a disciplined data channel, clearly classifying what can be inputted and selecting tools appropriate to the level of sensitivity.
Data risks when using AI in business do not come from AI automatically pulling your data. The risk comes from the habits of the team: pasting, working quickly, finishing tasks, without further thought. This is a real gap, and it expands proportionally with the number of people in the company using AI without common principles.
The most common behavior leading to data leaks is personnel pasting work content into the chat frame of public AI tools. It could be a contract that needs a quick summary, a piece of source code that needs debugging, a customer list that needs sorting, or a partner email that needs a reply. Each action may seem harmless on its own, but the data has left the company's control the moment it is sent.
It is noteworthy that many businesses tightly control emails, internal documents, and file access, but completely leave the AI chat framework open. This is a blind spot for oversight: a new channel, no rules yet, everyone handles it as they see fit.
A trustworthy AI system must protect data privacy and maintain resilience against attacks throughout the entire lifecycle of the system.
NIST AI Risk Management Framework (AI RMF)
Not every AI tool has the same level of commitment regarding data. The most important differences you need to understand:
Choosing the right tool for the right type of data is a management decision, not merely a technical one.
Acceptable with public tools: Draft communication content without confidential information, conduct market research from public data, brainstorm content ideas, and edit the style of texts without customer identification information.
Only use tools with security commitments or private environments: Summarizing or drafting contracts, handling customer data with names and contact information, debugging internal source code, analyzing strategic documents, any content that belongs to the company's intellectual property.
With customer data, Sinh Vũ maintains a strict boundary: sensitive information does not go to public AI tools. Not because technology is bad, but because this is part of professional caution and reputation. Saving a few minutes of operation is not worth the risk to customer trust.
What Sinh Vũ recommends to your team is to start from a simple internal principle: categorize your data into two groups, sensitive and non-sensitive, and then decide which tools are allowed to access each group. No need for complexity. Clarity is essential.
Regarding specific security configurations and legal constraints by industry, especially for businesses handling personal data under legal regulations, you should have data security and legal experts involved when establishing an internal AI framework. This is something Sinh Vũ cannot replace, and should not replace.
Topic: Data risks and security when using AI. Sinh Vũ guide, sinhvu.com
Select each item you find appropriate, then print or save as PDF to take with you.
If you have marked most of the signs above, this is the time to discuss in more detail. Sinh Vũ can help you review and propose a direction.
NIST AI Risk Management Framework (AI RMF); Press coverage on Samsung 2023 (Bloomberg, TechCrunch), summarized through AI security reports; GenAI Data Leakage: Employees Pasting Confidential Data into AI Tools (usecure). Insights on security configurations and specific legal constraints by industry are directional and require data security and legal experts to apply accurately.
It depends on the specific tools and packages used. Many free tools clearly state in their terms that data may be used to improve or train models. Paid packages for businesses often come with commitments not to use user data for training. You should read the data policy of the tool you are using before entering any customer or internal information.
Rewriting standard text is low risk, but the issue lies in the content of that text. If you copy a section from a contract, quote, or customer letter to edit the wording, sensitive data has already been exposed, even if the purpose is just to refine the style. The safe habit is to separate the content that needs editing from identifying or confidential information before putting it into the tool.