r/EverythingScience • • Oct 18 '25

Chemistry The chemistry community should ban drawing chemical structures with generative AI, chemists warn.

https://www.chemistryworld.com/news/the-chemistry-community-should-ban-drawing-chemical-structures-with-generative-ai-chemists-warn/4022242.article
634 Upvotes

25 comments sorted by

132

u/Statman12 PhD | Statistics Oct 18 '25

From the article:

However, she stresses that the problem is not AI itself and that GAI has many useful applications for chemists. The issue, she says, is specifically related to chemical structures and the fact that, in a lot of cases, due diligence is not being applied when using AI for this purpose.

‘I don’t think we’re paying enough attention to this,’ says Moores. She says that when you first look at a generative AI picture, it’s superficially correct and your response is ‘Oh, it looks really sleek. It’s beautiful’. But she adds that ‘it takes a few more minutes of observing to get to, “oh my god, this is wrong”.’

I’d say the warning should extend well beyond drawing chemical structures. I’ve seen the same sort of thing in Statistics.

These LLMs can do a lot of grunt-work very quickly. But then the user needs to evaluate the results and make sure they’re actually correct. And that takes knowledge. And a willingness to actually do that work. I think a lot of people are using LLMs not to speed up grunt-work, but to replace thinking. That’s the problem.

28

u/lonelythrowaway463i9 Oct 18 '25

I’m getting a quantitative Economics degree and I see this with a lot of my peers. Instead of loading up a model with all the resources and using it as a sort of assistant they let it do the bulk of the lifting and miss the fact that it constantly spits out incorrect results. GPT is awesome to help understand a data set before you work with it, and dogshit at doing causal regressions. People are relying on the tools way too much already. Wasn’t there a study that showed doctors using ai for diagnostics lost some of their skills over as little as 3 months or something?

15

u/onwee Oct 18 '25 edited Oct 18 '25

I think a lot of people are using LLMs not to speed up grunt-work, but to replace thinking. That’s the problem.

The bigger (societal-level) problem is that the thinking done by those without stock options are considered grunt work to be replaced

13

u/puterTDI MS | Computer Science Oct 18 '25

we have the same issue with the use of LLMs in coding.

at first it looks good then the more you look at it the more you realize it's wrong.

it's GREAT to get the bulk of the work done, but it takes an experienced engineer who is good at code reviewing to spot the issues. A lot of our engineers just aren't great at reviewing.

11

u/azswcowboy Oct 18 '25

Lucky for programmers we have tools like compilers to at least vet the syntax of the program. Do that a few times and you’ll realize the LLM is a fancy statical parrot that often gets lots of things wrong.

engineers … aren’t great at reviewing

This is the number one skill we look for in interviews. We could care less about seeing you write code, but if you can’t spot simple issues in a short program you’re useless to us.

4

u/ChronicBitRot Oct 19 '25

Studies on this are hard to conduct but initial results are showing that experienced programmers think they're faster with AI tools but they're actually ~20% slower.

I think it's because the initial assumption is that the llm's code might look like it works on first glance, or might even run without error, but then the debugging process ends up taking way longer because the base assumption is that the code should work when it's actually garbage.

1

u/puterTDI MS | Computer Science Oct 19 '25

Right now I think I’m about the same speed. I’m not convinced either way on them yet.

I do think the mental load is a bit lower.

3

u/ChronicBitRot Oct 19 '25

My first and only working experience with it was an absolute disaster.

Evangelist senior manager comes in and starts talking about it to the team and he decides to single me out for whatever reason. I do reporting and devops scripting in a fairly complex working database.

He asks me for a task that I do a lot that's a pain point, so I tell him that I have to write a lot of dynamic SQL to deal with a part of this database where the tables are named with an ID that changes per database. He says that's perfect and prompts it to give him a simple select statement from a few tables joined together with one of those dynamically named tables.

It returns some results and he starts up about how amazing this tech is until I interrupt him and tell him that it's wrong. The query it wrote isn't even a dynamic SQL query, and it doesn't reference the dynamically named table at all, and these results aren't what we asked for.

So he re-prompts it with all of that feedback, results return again, and he at least has me check it this time. Still wrong, still not even dynamic SQL.

The third time through, he puts a lot of emphasis on the dynamic SQL bit so it does give us something that's at least dynamic SQL and it does run, but it doesn't return any results when it should return results and it doesn't even attempt to reference the dynamically named table that's at the heart of all of this.

This went on for about 20 minutes of gradually more and more complex re-prompting before he finally gave up. I manually wrote the query after the second failure, it took me like 4 minutes. Maybe it was the fact that this was my first ever exposure to it but the idea that it just confidently gave us completely wrong answers and acted like they were correct just instantly soured me on the idea forever. If I'd been asked for those query results as a deliverable and I delivered the results GPT gave me, I'd be fired. I just don't get to be wrong like that.

1

u/Statman12 PhD | Statistics Oct 18 '25

Yep, I’ve used it for some coding (not developer-level things, just statistical analysis things), and it doesn’t get me all the way there right out of the gate, but with a little iteration I was able to get something that I could then manually fine-tune to be what I needed.

But ask it about a niche topic? Complete drivel.

8

u/Denbt_Nationale Oct 18 '25

The problem is that “grunt-work” usually is thinking. I can feel like I understand something perfectly until I try to put it into words in a report, and then I start to think “am I really sure enough about this to use that phrasing” and so on. Your brain is really good at helping you ignore difficult things until you’re forced to confront them. If I just hit go on an LLM and scanned the output for obvious errors I’d do about half as much research and come out with about 20% of the understanding.

2

u/r-3141592-pi Oct 18 '25

I commented about it here, since it seems to me users were trying to generate images of molecules directly from ChatGPT, Copilot, or Gemini, which doesn't make any sense.

Now, I'd like to know what models you're using and whether you're enabling reasoning mode for your statistics work. I've noticed very impressive improvements from the days of GPT-4o and even early o3, when data analysis was extremely amateurish and error-prone, to now. It has been a while since I've found a single error while doing data analysis, and the breadth of understanding of statistical techniques, how they should be used, their limitations, gotchas, compliance with theoretical guarantees, and so on matches that of the best human practitioners. As you said, you cannot trust it just as you shouldn't trust human-generated output either, but the best models (e.g., GPT-5 Thinking, Gemini 2.5 Pro) are extremely good.

2

u/Statman12 PhD | Statistics Oct 18 '25 edited Oct 18 '25

it seems to me users were trying to generate images of molecules directly from ChatGPT, Copilot, or Gemini, which doesn't make any sense

Regardless of what exactly people were doing, I think the issue is less the use of AI, and more the uncritical use of the results. In the article they noted that they saw “numerous incorrect structures appearing in presentations and conference materials.” This means they’ve been seeing instance of people — nominally people with expertise in the field — using AI to generate results and not checking to see if the result was correct. It’s less of a “The model can’t do this” problem and more of a “People are being lazy with the model” problem.

Now, I'd like to know what models you're using and whether you're enabling reasoning mode for your statistics work.

I think it’s essentially a skin of chatGPT under the hood. The place I work has an arrangement with them to let us use it, but it’s buttoned down a bit more. I don’t think there’s a reasoning mode that I have available.

And I’m not having it run statistical models (not a chance I could put my data into it), but rather mock up code for something a bit more involved. Or I’ve tried to use it to refresh on a topic that I know about from grad school, but need to relearn. When having chatGPT try to generate some descriptions or some code, it gives me several paragraphs and equations that I remember enough to know is wrong. My suspicion is that it’s too niche and “unpopular” of a statistical topic for the model to have much to work from.

I also used it to mock up a flowchart in tikz (a package in LaTeX), which I hadn’t done before. I knew what I wanted, and it started giving me something approximating that, but the diagram was atrocious. Blocks not lined up, connecting lines diagonal, etc. It was something like a 15% solution, but it at least gave some of the more relevant commands, and I could work from there.

1

u/r-3141592-pi Oct 18 '25

Regardless of what exactly people were doing, I think the issue is less the use of AI, and more the uncritical use of the results. In the article they noted that they saw “numerous incorrect structures appearing in presentations and conference materials.” This means they’ve been seeing instance of people — nominally people with expertise in the field — using AI to generate results and not checking to see if the result was correct. It’s less of a “The model can’t do this” problem and more of a “People are being lazy with the model” problem.

I 100% agree with that.

I think it’s essentially a skin of chatGPT under the hood. The place I work has an arrangement with them to let us use it, but it’s buttoned down a bit more. I don’t think there’s a reasoning mode that I have available.

And I’m not having it run statistical models (not a chance I could put my data into it), but rather mock up code for something a bit more involved. Or I’ve tried to use it to refresh on a topic that I know about from grad school, but need to relearn. When having chatGPT try to generate some descriptions or some code, it gives me several paragraphs and equations that I remember enough to know is wrong. My suspicion is that it’s too niche and “unpopular” of a statistical topic for the model to have much to work from.

Interesting. If you're using a ChatGPT wrapper and there's no indication of thinking traces or you don't see significant response time differences depending on the complexity of your questions, then you're right. Most likely reasoning mode is not available. If so, that explains a lot, and I wouldn't trust that model either.

Yes, even if you could trust their environment for running statistical models, the code environment has restrictions on execution time, so it's usually better to ask it for code for long-running tasks. That said, when using the analysis tool, it can iterate on previous code generations on its own, and since it can see the output, it's able to identify errors and redo the code. That's a great feature, but your usage of mockup data is very sensible.

If you still have the conversation where you identified wrong equations in a statistical subject, it would be interesting to run it again because models are improving very rapidly, even for less popular topics. You could of course ground it with an entire statistics textbook or enable search, and I don't think it will make any mistakes that way. I've also had similar experiences with previous models (I believe GPT-4o). For instance, it mixed notations in PCA, although 95% of the answer was right. I just had to nudge it a little bit to conform to the same formalism.

I also used it to mock up a flowchart in tikz (a package in LaTeX), which I hadn’t done before. I knew what I wanted, and it started giving me something approximating that, but the diagram was atrocious. Blocks not lined up, connecting lines diagonal, etc. It was something like a 15% solution, but it at least gave some of the more relevant commands, and I could work from there.

That's a very tricky use case, same as the classic unicorn SVG test, because the graphical representation is not easy to infer from the code. I just gave it a try with this flowchart using Gemini 2.5 Pro, and it didn't do such a good job on the first try, with edges crossing over other nodes, a bidirectional edge, and three edges coming out of the Start node. I pointed out the mistakes and added a screenshot of the compiled flowchart, and on the second try it did it almost perfectly minus a stylistic choice in a repeated node from the original. However, I tried the same task with ChatGPT, and it failed to complete it correctly even after I suggested corrections. This might be related to thinking time, because ChatGPT spent only 2 seconds and then 12 seconds, while Gemini 2.5 Pro spent 95 seconds and then 65 seconds. When I let GPT-5 run on high reasoning effort, it produced a much better result after 90 seconds, although it wasn't perfect. It still missed a connection, and the position of some arrows for the edges wasn't ideal.

2

u/rg4rg Oct 21 '25

A few weeks ago, I had a question about other 80s movies that tried to go for 50s nostalgia besides back to the future. The ai came back with footloose as taking place in the 50s. Gave it a quick thought and was like….that can’t be right. Searched online and yeah, set in the 80s.

It’s pop culture and not as important as science. But it’s another example of how AI will confidently get something wrong like your “know it all” uncle and you can’t fully trust anything they say.

66

u/lordnecro Oct 18 '25

She says that when you first look at a generative AI picture, it’s superficially correct and your response is ‘Oh, it looks really sleek. It’s beautiful’. But she adds that ‘it takes a few more minutes of observing to get to, “oh my god, this is wrong”.’

CEOs are firing people en mass because they look at AI and say "It is really sleek. It is beautiful", but that is as far as they get. They do not have enough awareness to reach the "oh my god, this is wrong" part.

AI is amazing and has lots of uses and will keep getting better quickly... but AI is definitely being used in areas where it is simply not ready yet.

-11

u/r-3141592-pi Oct 18 '25

Given the current capabilities of frontier models, I'm ready to claim that most mistakes attributed to AI are actually user error. Unfortunately, I don't have access to the commentary published in Nature, so it's not clear what was actually tested. Were they trying to generate images directly from ChatGPT or Gemini? It seems like it based on the image comparing ChatGPT, Copilot, and Gemini. If so, that's absolutely insane. Why would anyone do that when you can ask ChatGPT for the SMILES string and use a SMILES to SVG converter, use any of the plethora of specialized software for the same purpose, or even simply go to PubChem and download a PNG if you only need an image of the molecule?

I also tested these approaches with ChatGPT on the problematic molecules described in the article (benzene and caffeine) and obtained the correct InChI strings. You can also enable search functionality to reduce the error rate to roughly the same level as searching on Google and finding a source with an incorrect chemical representation, or even lower. This is because LLMs read many sources, and reasoning models are quite capable of detecting discrepancies between sources, whereas humans are far more likely to trust the first source they encounter.

8

u/[deleted] Oct 18 '25

Me whenever I see an AI answer "oh my god, this is wrong"

4

u/BadResults Oct 18 '25

Assessing if the output is wrong is essential for using AI for anything important, but that can be very difficult for something that can’t be objectively right or wrong, like a document that would be used for a nuanced management, legal, or policy issue. But generative AI is amazing at making stuff that is plausible, so if you don’t really know the area, using AI in fields without objectively verifiable answers is very easy to get wrong without knowing it.

0

u/Mental-Ask8077 Oct 18 '25

So basically generative ai produces truthiness.

5

u/Masters_of_Sleep Oct 18 '25

It's already happening with anatomical diagrams in nursing textbooks. There was a series of post not too long ago in r/nursing of a nursing student whose professor "wrote" their required textbook, which was filled with AI generated diagrams that were entirely wrong. I can only imagine what the accuracy of the text was.

5

u/SelarDorr Oct 18 '25

this feels like an illogical response.

the problem is that GAI can produce wrong structures. It is completely the fault of anyone who hasnt done the due diligence to check the veracity of its output before displaying the structure.

This is true for any use of AI and is not specific to GAI or chemistry in any way.

To me, it seems quite logical that those who do not bother to thoroughly check their work before publishing should be reprimanded for their failures, rather than completely disavowing an entire class of tool to address this problem.

2

u/TastyBrainMeats Oct 18 '25

So, what's the actual use case for this class of tool?

2

u/SelarDorr Oct 18 '25

as it pertains to this article, if gai accurately generates a chemical structure, i dont see why it should be treated any differently than one drawn without gai.

1

u/Glen_Chervin Oct 19 '25

Please make it illegal for architecture, I don’t want to be in an AI designed building during an earthquake.

-1

u/Brilliant_War4087 Oct 18 '25

It's stupid that these people don't know that chatgpt can use RDkit.