this post was submitted on 22 Apr 2024
33 points (97.1% liked)

Privacy

32665 readers
368 users here now

A place to discuss privacy and freedom in the digital world.

Privacy has become a very important issue in modern society, with companies and governments constantly abusing their power, more and more people are waking up to the importance of digital privacy.

In this community everyone is welcome to post links and discuss topics related to privacy.

Some Rules

Related communities

much thanks to @gary_host_laptop for the logo design :)

founded 5 years ago
MODERATORS
 

Hiya, just quickly wondering if anyone know about a good tool for comparing Privacy policies against each other? Im currently downloading each PP, then using self-hosted StirlingPDF to compare 1 on 1. However, I am looking for a more efficient tool, to compare multiple at the time, if there are any. Any tool that can handle multiple PDFs or HTML files and look at the differences between them kinda tool.

Appreciate any suggestions! ๐Ÿ•ต๏ธ

you are viewing a single comment's thread
view the rest of the comments
[โ€“] [email protected] 5 points 8 months ago (1 children)

Are you looking for a tool that can diff legal documents line by line or clause by clause? If the latter Iโ€™d bet an LLM with a large context size could do a pretty good job, especially if you used a script (or another pass through the LLM) to break them down into like sections so that could just compare e.g. all Controlling Law sections with each other and all IP Indemnification sections with each other.

Now that I think about it, tuning the prompt (and keeping the temperature very low, like 0) you could probably get it to return everything from proper diffs to summaries of conceptual differences. And it could definitely do multiples at once if you were to break them into like pieces ahead of time.

[โ€“] [email protected] 3 points 8 months ago (2 children)

Preferably line by line. Kind of like what Github does whenever you apply a commit, it will make a red line for what is removed and a green line for what is added code. I could look into LLMs though, but was hoping to find a quick n dirty tool to do the job.

[โ€“] [email protected] 4 points 8 months ago (1 children)
[โ€“] [email protected] 2 points 8 months ago (1 children)

This is pretty close to what im looking for actually, thanks for sharing! :)

[โ€“] [email protected] 2 points 8 months ago
[โ€“] [email protected] 3 points 8 months ago* (last edited 8 months ago) (1 children)

After reading this, I'm thinking whether converting the PDFs to markdown and diffing them with a text difftool could work.

If you go this route, you may want to test with different diff algorithms. Git has multiple too, but I don't remember right now which I found to be the best

[โ€“] [email protected] 1 points 8 months ago (1 children)
[โ€“] [email protected] 2 points 8 months ago* (last edited 8 months ago) (1 children)

Now that I'm at my computer, I was able to find the diff alg I was thinking about: it's histogram.

Here's an issue from gitea about when they changed the default git diff alg to this one: https://github.com/go-gitea/gitea/issues/23255
And here's an article I have found earlier about some of the available git diff algorithms, and when they are too be used: https://luppeng.wordpress.com/2020/10/10/when-to-use-each-of-the-git-diff-algorithms/

[โ€“] [email protected] 1 points 8 months ago

thanks very much for sharing โ˜˜๏ธ