Have you thought of active fingerprint pollution? Mathematically, find the search term that has the farthest embedding + noise than your current search term. You search for Harley Davidson, a parallel search is done for "3 mo old diapers".
Like another person said: passive fingerprint protection will be limited.
There is about a terabyte of XML IRS disclosures behind this number, so it's quite an exercise to derive. A few other interesting stats that explain the stat further:
- "Revenue” is not the same as donations - The the nonprofit dataset includes hospitals, universities, insurers and other tax-exempt businesses. Even patient fees, tuition, and investment income count as this revenue.
- The revenue skew is pretty extreme: $164k is the median, $6.43M is the average
- Scale attracts more scale: The concentration of revenue is increasing over the years
- As you can guess: hospitals and universities are among the largest orgs by revenue.
Definitely a different view of what a nonprofit means!
This is extremely problematic. Much of news is highly syndicated, so what looks like 10 credible sources are actually just 2. I think the best you can aim for is empirical sources: receipts, videos, photo evidence, public disclosures directly at the source. Archiving is ok if disclosed.
I don't think LLMs can discern truth when the majority of news sources are incentivized primarily for views.
Empirical sources - would be a first candidate to be implemented in a skill.
And you are right about 10 credible sources are actually 2. It will look like 10 independent sources to it, real gap. Skill does prefer primary source over secondary. For example it'll prefer NASA article over blog post who cited NASA article.
I'll work on empirical sources improvement in a next releases.
You may have overlooked this part of the parent's comment which addresses your syndication issue:
> Something along the lines of when you have two accounts that disagree traditionally and they agree on something. That's how you know it's likely to be true.
Outfits that blindly parrot talking points would not be at odds and thus not be good candidates.
It also doesn't help that many news sources will knowingly fudge the truth. There are topics about which they are constant sources of misinformation. Seems like depending on the topic (anything science related) you wouldn't want to use news sources at all. Other topics, you probably only want to use news sources. And arranging all of this weighting of sources is going to be difficult and controversial as well. Doesn't mean it can't be done. Doesn't mean it wouldn't be valuable. But there is a lot to it and most of it isn't so much about technology or math as it is about understanding who should and shouldn't be considered authoritative about what topics.
I can imagine this useful to 'compress' gigabytes of data onto a 2d canvas quickly. For that, I appreciate the effort.
One thing that would be useful is to read up on Ed Tufte's principles of data visualization. Many graph libraries don't implement basic visualization principles to make they key point clear, easy to see while still keeping the full depth and complexity of data visible.
After building a greppable dataset form 990 filings (~8.8M filings over 12 years), I've gotten more questions about how much leaders within nonprofits get paid. It turns out that the sector is extremely diverse: hospitals, insurance, grantmaking foundations, and hobby groups. The highest median top pay go to medical professionals rather than CEOs. There were also some interesting patterns across religious affiliation and geographic location of headquarters.
If there’s another way you’d like to see the data sliced, let me know and I’ll see what I can pull.
I'm also wondering if "perfect" and "good enough" are not really as important now vs. when a team of software engineers had to spend sprints implementing features. The rate of iteration is faster now, the rate of regenerating entire code bases is days. We can perhaps over engineer /more/ today than before.
The real challenge is reigning in the size and complexity of the codebase if you're using AI to generate it. I have a bunch of skills (YAGNI / KISS inspired) but it still requires a significant amount of human effort. Knocked down about 9000 lines of cruft this month and I know there's another 10k in there sitting around.
OP here: I visited the dialysis center in Jigjiga last year while doing field work with Amoud Foundation. The CKDu link to fluoride in groundwater was new to me. It seems most water infrastructure programs are scoped to access, not purification grade. Happy to answer questions on the data or the field visit.
That's an amazing contribution to literacy