Rendered at 10:35:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
syntaxing 6 days ago [-]
I like this so much. Maybe I’m romanticizing but hoping tools like this give local news a chance against monopolies like Sinclair Broadcasting.
cyanydeez 5 days ago [-]
It'll merely generate hyper local bots that generate hyper local slop. The media landscape is downhill.
ChrisMarshallNY 3 days ago [-]
I think it's just lead generation. If the paper decides to use AI to write/edit stories, that's another topic.
jasonlotito 3 days ago [-]
That's a long way of saying you didn't read the article.
binarymax 3 days ago [-]
Are search summaries slop?
ginko 3 days ago [-]
Yeah
binarymax 3 days ago [-]
It’s not a blanket ‘Yeah’. It’s a ‘sometimes’ and that sometimes depends on use. Distilling down a whole days worth of research is one of the best uses of AI.
irishcoffee 3 days ago [-]
My 12 year old told me that their friend group says "Oh I AI'd that" when they make a mistake, as a joke. Yet you trust AI to distill down a days worth of information in a 100% accurate way? When pre-teens already understand how unreliably they are?
binarymax 3 days ago [-]
If I'm getting citations, and the distillation has strong alignment with citations, then it's fine. In other words: the if the summary is a highlight and not a rewrite then I'm happy. I've studied this extensively and have even created metrics around this problem: https://maxirwin.com/articles/llm-rag/
irishcoffee 3 days ago [-]
So, if you review the entire dataset? At that point just write the summary yourself.
I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool.
Forgeties79 3 days ago [-]
> I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool.
I have been repeating some variation of this for months. If people would stop acting like it’s THE tech solution to ALL things ALL the time I bet a lot of critics would quiet down. The overhyping has become exhausting. It’s been going on for years.
GPT messed up the math for me the other day when I was simply adding 10 durations for a TRT. Couple of HH:MM:SS inputs, annoying to add up and I had it open.
It got it wrong, I told it it was wrong, it got it wrong again. Then out of curiosity I provided it the answer, asked for it to confirm it against the original numbers, and it went “you’re absolute right, it’s [original/wrong answer from earlier].”
This stuff happens probably 10-15% of the time for me regardless of the model. Not just math, just super simple crap. It’s wild to see at this point after 3-4 solid years of “hyperscaling” and overhyping. And it’s the kind of thing that keeps people like me from buying in beyond the foot or two we’ve stuck in the water.
ToucanLoucan 3 days ago [-]
I will say, while I agree with the broad strokes of what you're saying, in my experience when you provide an LLM with data to summarize, instead of it needing to go find it, the hallucination rate goes down staggeringly. I've never had a huge hallucination when the material I'm asking to be summarized is provided to it at the time of the request.
Any further distance than that though, even summarizing the conversation it's aware of up to that point, is dicier.
Zambyte 3 days ago [-]
What does "100% accurate" distillation even mean? That sounds contradictory. "Good enough" is fundamentally what makes distillation valuable, not perfection.
Conner_Hobbs 3 days ago [-]
[flagged]
phoghed 3 days ago [-]
Too many people use "slop" as a synonym for "AI Generated", so it's becoming useless as a term
karahime 2 days ago [-]
Because it has nothing to do with describing the quality of the thing and everything to do with trying to induce a flinch reaction.
Forgeties79 3 days ago [-]
If the work wasn’t so frequently sloppy we’d stop calling it slop.
phoghed 3 days ago [-]
Turns out slop is incredibly useful, valuable, and tasty
Forgeties79 3 days ago [-]
Lots of caveats and qualifiers needed here
1-6 3 days ago [-]
My local news station started focusing on issues far across the country to monetize on sentational headlines and clicks... I'm glad there's concern about improving local news.
ChrisMarshallNY 3 days ago [-]
Pretty cool.
I guess the downside might be, that someone could be replaced by this, but it seems that this is a perfectly valid application of AI.
In fact, lead generation seems to be a natural fit for LLMs.
mmooss 5 days ago [-]
> Early on, the team tried to balance quality with the high cost of searching many small, scattered sources. They discovered that prompting can implicitly control search depth and behavior, but only after trial and error.
I don't understand the cost here. The number of sources should be relatively tiny - it's not a general Internet search. The amount of data scraped would seem to be relatively tiny - how much data is there on suburban-county, PA?
mrweasel 3 days ago [-]
For the individual county it's probably not a ton of data, but multiply it up, it's a lot for all the counties in e.g. Philadelphia.
Normally there is data, meetings, events and so that you wouldn't report on in a newspaper, because it's really only relevant to maybe a few thousand people, but for those thousands it might be really important.
My city publishes a lot of stuff on their website, it's hard to find, relevant to maybe a thousand people, maybe less. The school board meetings, public utility companies, companies in general all publish massive amounts of information that's never surfaced, but is relevant to those living in the vicinity. Being able to collect all of this, sort it, assess it's relevans and produce hyper local news could be a massive boost for local grassroots movements and participation in local affairs and elections.
thephyber 3 days ago [-]
Agree that the permutation becomes a problem, but the organization should be super simple: the state mandates the use of some standard index file format for all municipalities governments, not unlike how sitemap.xml works for a website index.
There is a community-run project that standardizes each election's results across states, counties, and precincts. It could be a template for local info releases.
nemomarx 3 days ago [-]
do you have documentation on that standard index file? I'm not sure I've seen a Pennsylvania gov page explaining it
linkjuice4all 2 days ago [-]
The Philadelphia region accounts for roughly 2% of the U.S. population and often includes southern portions of New Jersey, northern areas of Delaware, and skirts portions of Maryland. Additionally the third circuit court and a variety of state and federal offices are located there so there are many unconnected systems with different owners.
We as nerds have failed to create a common standard and we have voters have failed to expand FOIA to surface all of the many meetings, documents, agreements, etc that govern our life so we're left with AI scraping tools that likely do a poor job to fill the gap.
tclancy 3 days ago [-]
I think it’s more about how many different types of content you’re looking at. Public google calendars, Facebook event pages, personal sites, etc.
4 days ago [-]
3 days ago [-]
FL410 3 days ago [-]
>The result: Scrape evolved from a small experiment to “load-bearing and critical infrastructure”
I use "AI" a lot, it is a fantastic tool. We need to stop pretending it is some kind of panacea. It's a tool.
I have been repeating some variation of this for months. If people would stop acting like it’s THE tech solution to ALL things ALL the time I bet a lot of critics would quiet down. The overhyping has become exhausting. It’s been going on for years.
GPT messed up the math for me the other day when I was simply adding 10 durations for a TRT. Couple of HH:MM:SS inputs, annoying to add up and I had it open.
It got it wrong, I told it it was wrong, it got it wrong again. Then out of curiosity I provided it the answer, asked for it to confirm it against the original numbers, and it went “you’re absolute right, it’s [original/wrong answer from earlier].”
This stuff happens probably 10-15% of the time for me regardless of the model. Not just math, just super simple crap. It’s wild to see at this point after 3-4 solid years of “hyperscaling” and overhyping. And it’s the kind of thing that keeps people like me from buying in beyond the foot or two we’ve stuck in the water.
Any further distance than that though, even summarizing the conversation it's aware of up to that point, is dicier.
I guess the downside might be, that someone could be replaced by this, but it seems that this is a perfectly valid application of AI.
In fact, lead generation seems to be a natural fit for LLMs.
I don't understand the cost here. The number of sources should be relatively tiny - it's not a general Internet search. The amount of data scraped would seem to be relatively tiny - how much data is there on suburban-county, PA?
Normally there is data, meetings, events and so that you wouldn't report on in a newspaper, because it's really only relevant to maybe a few thousand people, but for those thousands it might be really important.
My city publishes a lot of stuff on their website, it's hard to find, relevant to maybe a thousand people, maybe less. The school board meetings, public utility companies, companies in general all publish massive amounts of information that's never surfaced, but is relevant to those living in the vicinity. Being able to collect all of this, sort it, assess it's relevans and produce hyper local news could be a massive boost for local grassroots movements and participation in local affairs and elections.
There is a community-run project that standardizes each election's results across states, counties, and precincts. It could be a template for local info releases.
We as nerds have failed to create a common standard and we have voters have failed to expand FOIA to surface all of the many meetings, documents, agreements, etc that govern our life so we're left with AI scraping tools that likely do a poor job to fill the gap.
Please be satire