Skip to main content



🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


I saw a fireman releasing a bunch of female sheep on the side of the roads to eat all the brush in the hills. I asked him why they were all female.

He said, β€œOnly ewe can prevent forest fires!”

#funny #jokes #dadjokes

reshared this


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


If you write with #AI, you are accepting autocomplete text that a frequency-of-use search generated by a matching a machine-parsed request you made. You've turned yourself into an editor and a proofreader, the shit-jobs that go with writing; you aren't an author. Full stop. You don't learn more about the nuance of language by picking your words, you don't improve your style or delivery to talk more clearly or entertainingly to your reader, you don't grow. You atrophy. You deliver the lowest mean average of all the people's writing (as well as their opinions, misinformation, and bigotry) that the AI tool was trained† upon, and, increasingly, the average of #genAI text delivered by AI users the AI tool is also trained upon.

Guess what? Goes for all AI augmented professions where users abdicate using their creativity and experience to easily create mediocre content, choosing instead to become agents of quality control. Read this great article for a good synthesis of what I've been saying.

hackernoon.com/the-great-forge…

=.=.=.=.=.=
† Stolen is a better choice in context, though a much more loaded term than "trained upon." Many AI tools are trained on scraped data without royalties or asking permission. This includes your work in all likelihood.

#BoostingIsSharing

#Author #Writer #WritersOfMastodon #bookstodon #WritingCommunity #LLM #LLMs #ChatGPT

reshared this

in reply to RS, Author, Novelist, Prosaist

I forgot to add that if you are an #AI writer using #genAI to create your drafts, you've ceased doing the fun part, creating fictional whimsical worlds and captivating people other people wish to know, or synthesizing original ideas and bringing your original world-changing thought to readers. You're forfeiting being admired for your mind or your argument, instead you're bringing to others random fossilized thought scrounged from average ideas thrown together in an insipid soup of meaningless tokens masquerading as words mimicking thought. You've fed your ego a sugary breakfast cereal that's really only sawdust. You are sad.

#BoostingIsSharing

#Author #Writer #WritersOfMastodon #bookstodon #WritingCommunity #LLM #LLMs #ChatGPT
#genAI



In a way, I want to see ABC lose this, not that they don't deserve to win, but the only way they could lose that I can see is to declare that corporations don't have rights, and if they don't have rights, then they aren't people, right?




🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


#AskFedi: If I perchance wanted to quickly set up a hacky little website, and suggestions for a nice static site generator with some wider selection of templates (other features secondary for now, I can move later)?

#WebHosting #BoostsWelcome



in reply to Xoa Gray

@Xoa Gray I'd be curious, as I've done a pretty good job of staying under the radar, if'n I search for my name or image, little comes up because I use a lot of different names on the internet, not just OC names, but sometimes partial or complete fake names.
in reply to 🌴 Seph πŸ’­ πŸ‘Ύ

Names help a little, as well as spreading yourself across different platforms. But we've learned more recently that is a lot less of a countermeasure than it might have been before.






tumblr.com/rejectingrepublican…

#USPol #Trump



🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


Hey GNOME/GTK/Libadwaita nerds: is it even possible to create a (Bluecurve) theme that would work on Libadwaita applications? Like, an actual, real theme, not just a colour change or dark mode?

You people have no clue just how vague, contradictory, and confusing information about theming Libadwaita really is. I've heard people from the mentioned projects all say entirely different, opposing things about this.

I'm about to spend literally several euros per year on a stupid domain name for a stupid website to urge Fedora to bring back Bluecurve in the same way KDE brought back Oxygen/Air recently, and I'd like to get my facts straight - even if the endeavour is obviously going to be fruitless.




🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


πŸ”· Never be cruel. Never be cowardly. Hate is always foolish. Love is always wise. Always try to be nice, but never fail to be kind.

β€” The Twelfth Doctor, Twice Upon a Time

#DoctorWho

reshared this


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


tumblr.com/thenextdraft/825112…

#USPol #Hegseth #US-Mil


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


tumblr.com/dreamilypurpleoutca…






So, a follow up to that weird Hawaiin pizza I got…

Interesting, not sure what huli-huli sauce is, wasn't bad, wasn't great. But, I do really regret it, because I woke up with severe heart burn, breakfast turned my stop, and the heart burn has returned. I am so not enjoying this



🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


People have known for a while that AI companies are buying books, scanning them in, and destroying them in the process. Recently, unusually large anonymous purchases of "rare" books have been detected. AI companies at work? Now there's evidence!

"404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.

That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process."

404media.co/we-tracked-a-shipm…

(1/n)


We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility


Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.

A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.

That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.

β€œAmazon purchases books through commercial channels to help develop and improve the products and services our customers use,” an Amazon spokesperson told me in a statement.

The world’s AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing.

In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called β€œmodel collapse.”

πŸ“–
Do you know work at a facility where you scan books? I would love to hear from you. Using a non-work device, you can message me securely on Signal at @emanuel.404. Otherwise, send me an email at emanuel@404media.co.

These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.

In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.

404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment.

1

California


The shipment of books we tracked flew out of an airport in California.


2

Milwaukee


Next time we see it is right outside Milwaukee Mitchell International Airport.

1901 E College Ave, South Milwaukee, WI 53172

3

Distribution center


We then saw it in a warehouse belonging to a specialized shipping and distribution company Trifinity, right outside Kenosha Regional Airport, 30 miles south.

5312 104th Ave, Kenosha, WI 53144

4

Colorado Springs


It then starts heading West via truck. We see it on a highway outside Colorado Springs.


5

Truck Stop


It stops for the night at a trucking travel center.

2195 Hwy 6 and 50, Grand Junction, CO 81505.

6

Amazon warehouse: LAS8


The next day it arrives at an Amazon warehouse called LAS8.

5801 Nicco Way, Las Vegas, NV 89115.

The bookseller sent the shipment to Biblio, and it arrived at a California airport. It then traveled by plane to Milwaukee International Airport in Wisconsin. Later that day, the book traveled to a warehouse belonging to a specialized shipping and distribution company called Trifinity, right outside of Kenosha Regional Airport, about 30 miles south. The book remained there for about two weeks, at which point it began traveling west by what appeared to be a truck. I could see the book travel via the highway and spend a night over at a trucking travel center around Grand Junction, Colorado. The next day, the book arrived at its final destination, an Amazon warehouse called LAS8 in Las Vegas.

LAS8 is one of several large Amazon warehouses in the area, each with a different specialty. LAS7, right across the street, for example, is a fulfilment center, while LAS8 appears to mostly operate as one of Amazon’s β€œprint on demand” operations, which will print and ship books to Amazon shoppers as they are buying them. Initially, I was confused about why the book would arrive there, but Amazon employees who work at this location and who discuss working conditions there with other Amazon employees online explain that the the north end of the LAS8 warehouse, where I saw the book arrived, housed a different Amazon operation with a different code: VGT3.

Many Amazon warehouses have unique symbols to represent that specific site. VGT3’s symbol, painted on the entrance to the site and inside, is of a Tyrannosaurus rex, the massive carnivorous dinosaur, with its mouth open, holding an open book.
The entrance to Amazon's VGT3 facility.
β€œI work at VGT3 here in Vegas, and all we do is scan books,” one Amazon employee wrote on a forum for Amazon workers. β€œSome are assigned to cut books, and others go to receive where they get books and scan the bar codes. We didn't have rates, but now we do, but it's not stressful.”

β€œWorking at VGT3 is nice all we do is scan books,” another Amazon employee wrote. β€œIt's so cool.”

I saw VGT3 employees online talk about how working in this operation is a good but boring job because workers do the same, easy, repetitive tasks all day. I also saw some discussion indicating that VGT3 jobs are desirable for this reason, and because the warehouse sometimes offered night shifts. Earlier this year, some employees expressed concerns that Amazon would shut down the site because they had worked through the books they had and weren’t getting enough shipments of new books, though the site is still operational today.

That employees are scanning the barcodes or ISBNs on books β€” a unique serial number given to every published book β€” before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.

Elsewhere online in 2024, booksellers who sell their books on Amazon said they received a spike in orders to be shipped directly to the VGT3 location for a customer named β€œAmazon FC.” This was so unusual given its size they suspected the orders were some kind of scam. An Amazon representative chimed in to say they checked and that the orders were legitimate.

Like other major tech companies, Amazon is developing its own large language models (LLMs). Amazon considers its family of models branded Nova to be β€œfrontier” models, meaning it considers them to be competitive with other cutting edge LLMs from Google, OpenAI, and Anthropic. These models require massive amounts of training data in the form of human-written text. Internally, Amazon also used an AI coding agent called Kiro to develop its own software.

We first learned that AI companies wanted to scan millions of books for training data because of a lawsuit from book authors against Anthropic, which revealed Anthropic’s β€œProject Panama.” The goal of the project was to acquire books from commercial bookselling marketplaces, cut the spines off the books, and scan them. It’s possible to scan books without destroying them, but cutting the spine makes it cheaper and faster. Additionally, the judge in the lawsuit ruled it was fair use and not a copyright violation for Anthropic to scan a book for training data in part because it destroyed the original, printed copy. Essentially, it’s the customer’s right to take physical media and store it digitally, and destroying the original copy means that copy isn’t duplicated and resold, and isn’t cutting into the publisher’s business.

We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak. As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.

β€œThere are different types of value,” the bookseller said. β€œThere's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together.”

Amazon provided its statement above but did not respond to questions about why it cut the books it scanned, how many facilities like this it has across the world, and their response to the backlash from people upset that AI companies are destroying books.


This entry was edited (1 day ago)

reshared this

in reply to John Carlos Baez

Anthropic has been seeking to "destructively scan all the books in the world" since at least 2024, according to the Washington Post:

"In early 2024, executives at artificial intelligence start-up Anthropic ramped up an ambitious project they sought to keep quiet. β€œProject Panama is our effort to destructively scan all the books in the world,” an internal planning document unsealed in legal filings last week said. β€œWe don’t want it to be known that we are working on this.”

Within about a year, according to the filings, the company had spent tens of millions of dollars to acquire and slice the spines off millions of books, before scanning their pages to feed more knowledge into the AI models behind products such as its popular chatbot, Claude."

washingtonpost.com/technology/…

(2/n)

This entry was edited (1 day ago)
in reply to John Carlos Baez

More recently, antiquarian booksellers have been seeing suspicious increases in sales. And the reason is pretty clear:

"Now, as 404 Media reports, this practice has become prevalent enough that even well-established book sellers are looking to cash in on the AI boom. One called ISBNdb, which boasts the β€œworld’s largest book database,” extolls that the β€œworld’s best AI training data is setting on a shelf,” upholding these physical texts as uncorrupted by shoddy AI writing that’s already polluted so much of the internet (and indeed, newer books).

[...]

Once focused on helping libraries, distributors, and book shops find and sell books, ISBNdb now helps AI companies bulk-buy anywhere between 1,000 to one million books per order, according to 404.

As an added bonus, it also promises AI companies that it’ll keep their purchases under wraps β€” nobody wants to end up in the spotlight like Anthropic and Meta, obviously β€” while clearly sounding aware about how incredibly shady the practice sounds.

β€œThe optics problem is real,” ISBNdb’s site says. β€œβ€˜AI company destroys two million books’ is not a headline that generates sympathy.”

One small book seller said that in April, he suddenly went from selling no more than 20 books a week to hundreds, and he’s almost certain that the customers are AI labs, noting the random selection of the books and how they all have ISBNs. He added that his inventory is full with rare and out of print books, meaning that an AI company could be destroying some of the few remaining copies that can be found."

futurism.com/artificial-intell…

(3/n)

This entry was edited (1 day ago)
in reply to John Carlos Baez

Kinda disappointed to see a science communicator uncritically repeating nonsense claims like "upholding these physical texts as uncorrupted by shoddy AI writing that’s already polluted so much of the internet (and indeed, newer books)." Outlets like 404 Media were pushing this misconception earlier in the year, that AI-generated text is somehow harmful for training and will lead to a situation where models get worse because they're training on increasingly "sloppy" writing over time. In fact LLMs are trained predominantly on LLM-generated text, and it's one of the reasons performance has improved so fast over the last year.
in reply to Tufty Indigo πŸͺ—

@tuftyindigo
Where does this idea come from: β€œIn fact LLMs are trained predominantly on LLM-generated text, and it's one of the reasons performance has improved so fast over the last year.”?

Can you point me to sources?
Thanks in anticipation. πŸ˜€πŸ™πŸ»

#LLMs #LLMGeneratedText #LLMPerformance

@johncarlosbaez

in reply to Su_G

@Su_G
arxiv.org/abs/2503.03150 is about why "model collapse" doesn't work like that, and arxiv.org/abs/2510.01631 out of Meta FAIR is an empirical explanation of what training processes prevent model collapse. Going back to the fundamental question about synthetic data, arxiv.org/abs/2501.12948 describes how distillation (using a stronger model to generate training data for a weaker model) allows models to perform better than using natural training data alone.
@Su_G


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


What a fab idea. In Brazil prisoners earn time off their jail sentence by reviewing books.


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


bgr.com/2233665/solar-farms-un…

Why do you need any grading for solar panels, aren't they post mounted?

#Green-future #Agrivoltage #Solar #Solar-Panels



🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


huffpost.com/entry/picasso-etc…

Kind of want to know how it wound up there


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


So, we're rocking a brand new trackball after the last one died, one that ran me a whopping $30 from Amazon.
Yeah, its what you'd expect, alright but nothing special. I did have some trouble with it being too sensitive at first, but I just tweaked that in the mouse settings. Not sure if'n its the problem, or the old trackball I was limping along just wasn't moving quickly any more. That and I've been using a cheap mouse for a week or so, so that might've altered what I expect.
Either way, it works, I miss the height adjustability of the Perimice one I had before, but this works. I mean its basically just a generic Chinese mouse flipped upside with a ball and rollers added. Really the only question at this point is how long will it last?

#Trackball #Computers #Mice #Pointing-Device #Electronics







Peter on Mastodon
"The Flock Sock is a slip on cover designed to help protect a Flock camera from the harmful radiation of the sun. Your city spent their very own hard earned tax dollars scattering these cameras everywhere to keep you safe, all it cost you was a little privacy. The very least you can do is help protect these valuable civil servants from getting a bad sunburn."
Makerworld



Chuff² reshared this.


🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


So, this week's pizza is the weird Hawaiian Pizza I got at Aldis. We all know what's on a Hawaiian, right? Pineapple and Canadian bacon, although sometimes ham or sausage replaces the Canadian bacon.
This one? Chicken, pineapple, green peppers and huli-huli sauce. Well this'll be interesting

#Pizza #Hawaiin




🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


tumblr.com/whitepeopletwitter/…

🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


Trump Orders the USS Abraham Lincoln to Deploy to Ford's Theater Next #ABlueView 🧡
1/2

🌴 Seph πŸ’­ πŸ‘Ύ reshared this.


It is the year 2036.

You wake to discover your subscription to ClaudeFS has expired, locking you out of all of your storage. You try to pay for another week but the chatbot at the bank read your shitposts last night and is refusing to serve you for security reasons, because your vibe doesn't match your social media presence.

All you wanted to do was a little bit of drawing, but ever since it became illegal for anyone but LLM companies to store data outside the cloud, you don't want to risk pulling your WIP backup in case you get caught with downloaded data.

reshared this

⇧