13 Comments
User's avatar
Sonya Sanford's avatar

This is so well explained (and shown), and such an important conversation to be having!

The Soul Food Stack's avatar

Thank you from the creator community. We share our cuisine, culture, and stories with so much heart and soul! ♥️

Magdalini Zografou's avatar

I have been following your blog for so long, ever since before social media. I have had my food blog since 2009 myself. I am so happy to have found you here on Substack via Diane Jacobs' newsletter.

AI is so bothersome but I'm afraid it's here to stay. Most know, or suspect, that something is AI but don't care. Thanks for offering information and ways for people who do care to do something when they spot content theft etc. It's so important for all creators.

Yvette Marquez-Sharpnack's avatar

Hola! So lovely to connect here on Substack. Thank you for following my cooking journey.

Mao Zhou's avatar

Is AI really the problem or plagiarism ?

Yvette Marquez-Sharpnack's avatar

It certainly is a bit of both - AI and plagiarism. When these type of graphics then lead to a site that is not human -- that is AI generated.

N G's avatar

Thank you for spreading the word and for including my examples!

Betty Williams's avatar

Such good examples, Yvette! And I'm so happy that you are speaking out about this. Did you see that Bjork and Lindsay of Pinch of Yum had their entire site ripped off and duplicated, including AI manipulated photos of them and their children? It was sickening!

Yvette Marquez-Sharpnack's avatar

Yes, I heard about it and I saw it on the Bloomberg article that I was also featured on. My blog at one point was also scraped, but they didn’t even bother to manipulate any photos. They just used everything and made it look like it was completely their site. Thankfully, somebody shared it with me and I was able to have it taken down within a week. But it is sickening that people do this.

Theresa's avatar

This was a well written article. I agree with the content and the fact that I see AI as a great tool for science and technology issues, but to be able to rip off everyday people who are earning a living creating content wether it be recipes, journalism or any other medium created by a person and then stolen without permission of any kind is wrong.

Yvette Marquez-Sharpnack's avatar

Thanks for sharing this perspective. It’s definitely a frustrating issue for creators, and anything that helps reinforce ownership and source of truth is worth paying attention to. Appreciate you adding to the conversation.

Ted Fay's avatar

It’s a challenging process, and frustrating to all creators. Scrapers using AI to "spin" your creative and intellectual property (and family history!) and repost them. Fighting this by blocking scrapers is like playing Whac-A-Mole, with Google even working hard to fight these technologies.

In addition to your steps, there are some more modern, "defensive publishing" approach that might interest you.

Instead of “hiding” content by blocking engines, or trying to block the door, give it a Digital Birth Certificate. By enhancing the "hidden" code (Schema) that your site already sends to Google and the like, you can make your build the technical ownership of your content, enhance its value to Google, Bing, and other crawlers,and build a landscape to work to address scrapers.

• Verifiable Provenance: Linking the recipe directly to your physical cookbooks and family history in a way scrapers will struggle to replicate. Not just in the visible content, but deeply in the web page’s schema. This will also help build the authority and trustworthiness of your content, which is what the “good” crawlers from Google, etc. are seeking.

• Temporal Proof: Using "digital notaries" to prove you published the content first, months or years before a scraper, and embed this proof on your content. Not just an “article published on” type listing, but leveraging known 3rd parties. Builds authority of your content as well.

• Ownership Signals: Embedding machine-readable licensing and SHA-256 content hashes that tell Google's algorithm exactly who the "Source of Truth" is.

Essentially, when a scraper takes your content, they’d be taking your "ID card" with it, making it much easier for search engines to identify them as the duplicate and you as the authority.

None of this is foolproof, nor is any of it instant. But you can start with the tools that your Wordpress site likely already has.

Helena-Laura's avatar

This was such great insight once again! And I have been wondering if I'm really the only one who is paranoid and suspicious about some weird accounts on Facebook and endless recipe lines, which come in at around 10 recipes per day or even more. The photos are soulless, and the captions are pointless. But people like, comment, and share like crazy! It's really disheartening! I guess they're even here in Substack now, at least it seems like it, when exploring some accounts. Thank you for seeing all this and sharing your thoughts!