Model Values Become a Distinct AI-Development Challenge
As open-source models approach closed systems in capabilities and economically valuable tasks, the values embedded in models are becoming a separate development challenge. Values, ethics and morality are shaped—implicitly or explicitly—through pre-training, mid-training, post-training, classifiers and safety systems.
The unclear values of Chinese open-source models are cited as one reason many American companies are reluctant to use them. A key unresolved research question is whether model values can be organically shaped for a country, company or individual, including through possible “post-post-training.”
Trust in benchmarks is described as being at an all-time low and potentially declining further, complicating how the value of new models should be communicated.
