• Hacker News
  • new|
  • comments|
  • show|
  • ask|
  • jobs|
  • simontwiggyahuk 3 hours

    https://discuss.google.dev/t/trusted-automation-with-google-...

  • bob1029 6 hours

    I have presented approaches like this to banking clients and they are still not very interested. The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.

    If your product is handling natural language conversations from end customers, there is not much you can do to prevent the occasional PII leak without ruining the rest of the pie. ZDR is your best mitigation if you actually want the magical AI experience to work the way the investors hope it can.

    PII can often become disclosed by way of many correlated factors that are not considered PII on their own. Even a perfect AI system cannot capture all of these relationships. You could probably locate where I live within a 20 mile radius if you spent enough time analyzing my HN comments over the years. Not one of these comments on their own would trigger a PII filter.

  • nhinck2 6 hours

    98.4% is nowhere near good enough to call it PII redaction.

  • iAMkenough 6 hours

    So that’s why big ballz (co author of the OP) put our PII in an insecure AWS instance via DOGE’s starlink terminal

    https://www.csoonline.com/article/4046997/whistleblower-doge...

  • handfuloflight 5 hours

    Why did the National Design Studio see the need to put all the readable text on the right column of the page?

  • theritik12ee1 54 minutes

    [flagged]

  • swiftcoder 5 hours

    What we'd really like is a PII redaction model for video...

  • throw03172019 1 hours

    Plain text in a chat input is only one piece of the problem. What about files like PDFs and documents filled with PII.

  • dwa3592 6 hours

    I have worked in this field and I am the author of this package - https://github.com/deepanwadhwa/zink

    A few things jump out since this is done by the government:

    - the lowest hanging fruit for this problem is to clearly tell people (citizens) not to share any personal info with chatbots which can cause financial harm or identity theft. the example on the page shows a person sharing their SNN with a chatbot to help them find an apartment - "My name is Maria Garcia, my Social Security number is 123-45-6789, and I make $1,950 a month. Can you help me find affordable housing?" - why?? this is the opposite of what i would expect a government to advise their citizens.

    - it's never too late for a good policy; the government should have extended HIPPA and other data privacy laws to AI companies - the AI company must not store anyone's SSN, no matter how stupid the user is. It should be on the AI company to not store it; so this type of layer should be on the AI company's side.

    - technical; there are quasi identifiers of privacy (that's what my package targets) that are asymptotically hard to to deal with - meaning - if you remove everything that can leak your privacy the text would become meaningless. i don't think rampart can solve for that either and it should be clearly said on the website.

    bberenberg 5 hours

    I think all of your points are very valid, but it doesn’t remove the point that the government is trying to make it free and easy for application builders to do a little bit better than they are today. This is commendable on its own, even if it’s not perfect.

    applesauce3572 2 hours

    Point is completely fair. It seems to be sort of a grey area that they skim over entirely on your point about sharing deep personal info.

  • sgnelson 6 hours

    The skeptic in me really can't trust our current government to protect my information.

    Cider9986 5 hours

    You could never trust the US government to protect our information since it moved online.

    BowBun 6 hours

    Good thing you don't need much trust in this case. Source available here - https://github.com/nationaldesignstudio/rampart

    I suppose the model could be doing stuff, but you can also switch that out for your own with the source.

    goodmythical 6 hours

    I mean, it can be run locally, so you don't necessarily have to trust it, unless there's been any model-as-a-vector CVE.

    That hasn't happened yet, has it? Where running a model from HF directly compromises the machine as opposed to some breakout or exfil done by the model after the fact?

    That said, if you're filing your taxes or have a 'real' ID, the government already has all pertinent information required to fuck you over either deliberately or via a leak, so...

    Ohentis 2 hours

    The vulnerability to watch out for would be the model being trained to not redact some specific information.

  • Cider9986 5 hours

    The source code should be public domain, no?

    22c 5 hours

    https://github.com/nationaldesignstudio/rampart ?

    iAMkenough 3 hours

    LICENSE errors out https://github.com/nationaldesignstudio/rampart/blob/master/...

    CC by 4.0 is not the same as CC0 (public domain): https://creativecommons.org/public-domain/

  • Onavo 4 hours

    Is this good enough for HIPAA?

    fwip 3 hours

    No.

    Onavo 1 hours

    It's the technique used by a lot of AI healthtech

    estearum 1 hours

    What is? This is nowhere close to what HIPAA requires.

    Onavo 35 minutes

    A lot of clinical decision support AI tools essentially scrub the PHI data on the frontend.

    estearum 29 minutes

    PHI is a different set of data points from PII amigo.

    Onavo 22 minutes

    Same technique is good enough.