You are currently viewing Do AI Systems Have Scholarly Cultures?

Do AI Systems Have Scholarly Cultures?

The more I work with LLMs in my research, the more it seems to me that we need to think more carefully about what exactly we are interacting with when we use something like ChatGPT, Claude, DeepSeek, etc. We tend to talk about “models,” and certainly the model is extremely important, but increasingly I think that focusing simply on the model obscures quite a bit of what is actually happening.

A model is the trained neural network itself. It has learned statistical patterns from enormous amounts of data and uses those patterns to predict and generate text. However, when we use ChatGPT or Claude or some other AI system, we are not interacting with a bare model. We are interacting with a larger system that has been built around that model, and one important part of that larger system is what people increasingly call a “harness.”

The harness is basically the software surrounding the model that helps determine how we interact with it and what it can do.

When ChatGPT first appeared, from the perspective of the user, that harness was relatively simple. It managed the conversation, kept track of some context, and passed prompts back and forth to the model.

Now, however, a harness can connect a model to tools, allow it to take actions, manage memory, plan and sequence tasks, check its own work, coordinate multiple models or agents, search the Internet, read files, execute code, and so on.

That said, the distinction between “model” and “harness” is itself a simplification, because there is another important layer in between: post-training. A model is first trained on a huge amount of data, but it is then further trained to follow instructions, behave in certain ways, avoid certain kinds of responses, prefer certain kinds of answers, etc.

So, when we interact with an AI system, what we encounter is really the result of multiple things working together: what the model learned in its original training, how it was subsequently trained to behave, and how the harness directs and equips it when we actually use it.

From the outside, of course, we usually have no way of knowing exactly which of these layers is responsible for a certain behavior. Nonetheless, the more I use different AI systems, the more I get the sense that the differences between them are not always simply differences in what they “know.” Instead, they often seem to differ in what they pay attention to, what they think needs to be checked, what they assume is important, and what they regard as a satisfactory answer.

Recently, for instance, I have been working quite a bit with DeepSeek. It is great (and very fast) when it comes to things like optical character recognition of printed Chinese texts, and it is also very good at translating those texts. However, when I ask it to examine those same texts in the context of historical research, I sometimes find that I have to keep pushing it in directions that I would have expected it to go on its own.

By contrast, I find Claude and ChatGPT much better at doing that. Again, I do not mean that DeepSeek does not “know” the relevant information. In many cases, it clearly does. What seems different is that Claude and ChatGPT are more likely to anticipate certain kinds of scholarly issues without my explicitly telling them that those issues matter.

This is where things start to get interesting for me, because in the case of my work on Vietnamese and Southeast Asian history, the difference between working with something like Claude and working with DeepSeek sometimes feels a bit like the difference between working with a good Western scholar and a good Chinese scholar.

Scholars working in different academic traditions operate according to somewhat different conventions, and over time they develop different scholarly instincts. They know what their readers will expect to see explained. They know what kinds of evidence will need to be cited. They know what terminology should be retained in the original language, what needs to be Romanized, what needs to be translated, and so on.

A scholar can therefore possess all of the relevant information and still produce something that looks different to someone working in another academic tradition. The issue is not really knowledge. It is knowing what needs attention.

I recently encountered a small but very clear example of this. In the previous post, I shared a translation of Taiwanese scholar Chen Ching-ho’s “On the ‘Hạ-châu Missions’ (下洲公務) Conducted during the Early Period of the Nguyễn Dynasty.” The article was originally written in Japanese and was then translated into French. I wanted to produce a combined English translation that would employ the Romanizations used in the French text while also restoring the original Chinese characters for key terms and cited passages that appeared in the Japanese text.

That creates a fairly complicated editorial problem for a work in English. A Chinese scholar writing in Chinese generally does not have to worry much about Romanizing Chinese characters. The characters themselves are right there. However, in an English-language article on premodern Vietnamese history, Romanization can become surprisingly complicated.

A Chinese term may need to appear in Pinyin in one context, while the same characters may need to be read in Sino-Vietnamese in another. Vietnamese names need Vietnamese spelling. Chinese names need Chinese Romanization. Certain historical terms may need to appear both in Chinese characters and in a Romanized form. Then there are older French Romanizations, Vietnamese readings of Chinese names, Chinese readings of Vietnamese historical terms, etc. None of those things are particularly difficult by themselves, but you have to know which convention belongs in which context.

DeepSeek clearly knows all of this. It can convert Chinese characters into Pinyin without difficulty. It can give Sino-Vietnamese readings. It understands the texts. However, when I asked it to work on this translation, it kept struggling to decide what needed to be Romanized and in what way. I would correct it, explain what I wanted, but each time I would discover something that was not the way it should be.

Eventually I gave up and asked Claude to do it, and—voilà!—it immediately understood what I wanted, even though I gave it relatively limited instructions.

What I found interesting was that this did not seem to be a case where Claude “knew” something that DeepSeek did not. DeepSeek clearly possessed the relevant knowledge. Rather, Claude seemed more inclined to recognize that I was producing something that was supposed to look like an English-language academic article on premodern Vietnamese history and therefore to anticipate certain conventions associated with that kind of scholarship.

In other words, the difference was not necessarily one of knowledge. It was a difference in what the system regarded as important.

And this, I think, is something that historians will understand quite easily. A historian trained in one academic tradition will produce work that takes a different form from a historian in another academic tradition.

The more I work with different AI systems, the more I wonder whether something similar is happening there as well. Again, I do not think that we can simply say, “This is the harness.” The behavior we encounter probably comes from a combination of the original training data, post-training, system instructions, tool use, memory, and the harness that coordinates all of those things.

While we usually cannot see enough of that process to determine exactly where a particular tendency comes from, all of those layers (for now) involve human choices. People decide what material a model is trained on. People decide what kinds of answers should be rewarded during post-training. People decide what instructions the system receives. People decide what tools it can use and what kinds of things it should check before answering.

Those decisions inevitably contain assumptions about what matters. So, would it really be surprising if different AI systems developed different “defaults”? And if they do, would it be surprising if some of those defaults reflected the intellectual or cultural environments in which those systems were developed?

Perhaps “culture” is too strong a word. However, I do increasingly feel that different AI systems have different scholarly instincts, and that those habits can sometimes resemble differences that we see between scholarly traditions.

It might be the case that this is especially noticeable in the humanities because so much of what we do cannot be reduced to simply getting the “right answer.” In computer code, for example, there is often a fairly clear way to determine whether something works. You run the code. If it breaks, something is wrong. Historical scholarship is not like that. It depends heavily on judgment, convention, context, terminology, citation practices, and the ability to recognize what needs to be checked.

And that may be precisely where these differences between AI systems become most visible.

So perhaps in working with AI systems the important question for historians is not simply, “What does this model know?” Instead, it is, “What has this system been trained, instructed, and equipped to notice?”

Finally, while the example I gave is one that relates to editorial-level issues, I can see the same differences between models at a deeper level as well. It is just that the editorial-level example is easier to explain and visualize.

Subscribe
Notify of
guest

0 Comments
Oldest
Newest Most Voted