Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> *There is always a tradeoff between privacy and utility. The only way to achieve 100% private data is 100% noise, but the privacy-utility tradeoff curve isn't linear, and you can still achieve very good utility and very good privacy in many cases, especially with the best tools. Methods are also improving over time, reducing the impact of the tradeoff.

This is the crux of the problem. In every situation where I've seen it deployed, "anonymization" is as far up the utility curve as the corporate interest can get while still credibly calling their data anonymized. The standards and rules that are in place for restricting what types of data can be called "anonymized" to the public are so far up the utility curve that there's functionally no privacy at all.

I don't know what your startup does or what steps it takes to make that tradeoff responsibly, but assuming that you are responsible, you are one of the few, and most of the customers in the world would rather use a client that is using the word "anonymous" more irresponsibly because:

1. they can get away with it

2. there's a lot more business value in obfuscating data as little as possible

The EFF article is bang-on here, because the majority of the "anonymization" in the space isn't being done by responsible folks (like yourself, presumably), it's being done by psychopathic business interests.



My experience is no different and in fact it’s hard to convince anyone to treat indirect identifiers at all. Everyone wants the bare minimum, which is unfortunate, but understandable for someone trying to run a business. All of our current customers are only using their data internally and anonymizing to protect identities as that data is shared across other internal groups, mostly for testing and analysis. Even protecting only direct identifiers is far more responsible than the vast majority of businesses, though we push for indirect treatment.

If you treat indirect identifiers with our system, the minimum level of anonymization would ensure that there are always a minimum of two matching combinations of indirect identifiers, i.e. k=2, for every record in the dataset. This is far from foolproof for preventing data leakage, but it does prevent singling out and provides a lot of protection in the case of breach.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: