EXAONE 3.0 7.8B Instruction Tuned Language Model
- URL: http://arxiv.org/abs/2408.03541v3
- Date: Tue, 13 Aug 2024 10:09:32 GMT
- Title: EXAONE 3.0 7.8B Instruction Tuned Language Model
- Authors: LG AI Research, :, Soyoung An, Kyunghoon Bae, Eunbi Choi, Stanley Jungkyu Choi, Yemuk Choi, Seokhee Hong, Yeonjung Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Yountae Jung, Euisoon Kim, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Youchul Kim, Edward Hwayoung Lee, Haeju Lee, Honglak Lee, Jinsik Lee, Kyungmin Lee, Moontae Lee, Seungjun Lee, Woohyung Lim, Sangha Park, Sooyoun Park, Yongmin Park, Boseong Seo, Sihoon Yang, Heuiyeen Yeen, Kyungjae Yoo, Hyeongu Yun,
- Abstract summary: EXAONE 3.0 instruction-tuned language model is the first open model in the family of Large Language Models (LLMs)
EXAONE 3.0 demonstrates highly competitive real-world performance with instruction-following capability against other state-of-the-art open models of similar size.
Our comparative analysis shows that EXAONE 3.0 excels particularly in Korean, while achieving compelling performance across general tasks and complex reasoning.
- Score: 41.95996640625627
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: We introduce EXAONE 3.0 instruction-tuned language model, the first open model in the family of Large Language Models (LLMs) developed by LG AI Research. Among different model sizes, we publicly release the 7.8B instruction-tuned model to promote open research and innovations. Through extensive evaluations across a wide range of public and in-house benchmarks, EXAONE 3.0 demonstrates highly competitive real-world performance with instruction-following capability against other state-of-the-art open models of similar size. Our comparative analysis shows that EXAONE 3.0 excels particularly in Korean, while achieving compelling performance across general tasks and complex reasoning. With its strong real-world effectiveness and bilingual proficiency, we hope that EXAONE keeps contributing to advancements in Expert AI. Our EXAONE 3.0 instruction-tuned model is available at https://huggingface.co/LGAI-EXAONE/EXAONE-3.0-7.8B-Instruct
Related papers
- K-EXAONE Technical Report [76.23621600385238]
K-EXAONE is a large-scale multilingual language model developed by LG AI Research.<n>It supports a 256K-token context window and covers six languages: Korean, English, Spanish, German, Japanese, and Vietnamese.<n>We evaluate K-EXAONE on a comprehensive benchmark suite spanning reasoning, agentic, general, Korean, and multilingual abilities.
arXiv Detail & Related papers (2026-01-05T02:30:59Z) - EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes [42.31740630042654]
EXAONE 4.0 integrates a Non-reasoning mode and a Reasoning mode to achieve both the excellent usability of EXAONE 3.5 and the advanced reasoning abilities of EXAONE Deep.<n>The EXAONE 4.0 model series consists of two sizes: a mid-size 32B model optimized for high performance, and a small-size 1.2B model designed for on-device applications.
arXiv Detail & Related papers (2025-07-15T15:24:51Z) - Command A: An Enterprise-Ready Large Language Model [180.18356391290172]
Command A is an agent-optimised and multilingual-capable model.
It offers best-in-class Retrieval Augmented Generation capabilities.
arXiv Detail & Related papers (2025-04-01T12:08:07Z) - EXAONE Deep: Reasoning Enhanced Language Models [35.326172288018505]
We present EXAONE Deep series, which exhibits superior capabilities in various reasoning tasks.
We train our models mainly on the reasoning-specialized dataset that incorporates long streams of thought processes.
arXiv Detail & Related papers (2025-03-16T14:39:33Z) - EXAONE 3.5: Series of Large Language Models for Real-world Use Cases [35.04562823885241]
The EXAONE 3.5 language models are offered in three configurations: 32B, 7.8B, and 2.4B.
For commercial use, please reach out to the official contact point of LG AI Research.
arXiv Detail & Related papers (2024-12-06T08:53:46Z) - Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier [72.5652085347547]
We introduce the Aya Expanse model family, a new generation of 8B and 32B parameter multilingual language models.
By leveraging several years of research at Cohere For AI and Cohere, Aya Expanse sets a new state-of-the-art in multilingual performance.
Our evaluations on the Arena-Hard-Auto dataset, translated into 23 languages, demonstrate that Aya Expanse 8B and 32B outperform leading open-weight models.
arXiv Detail & Related papers (2024-12-05T15:41:06Z) - \llinstruct: An Instruction-tuned model for English Language Proficiency Assessments [6.307485015636125]
We present an 8B instruction-tuned model that is designed to generate content for English Language Assessments (ELPA)
Our work involves creating a new dataset of 70K instructions and explanations in the ELPA domain.
Human evaluations are conducted over unseen instructions to compare these SFT models against SOTA models.
arXiv Detail & Related papers (2024-10-12T00:47:45Z) - Rethinking Optimization and Architecture for Tiny Language Models [39.892066839422796]
The application of language models on mobile devices is facing huge challenge on the computation and memory costs.
In this study, based on a tiny language model with 1B parameters, we carefully design a series of empirical study to analyze the effect of each component.
Several design formulas are empirically proved especially effective for tiny language models.
arXiv Detail & Related papers (2024-02-05T07:59:38Z) - YAYI 2: Multilingual Open-Source Large Language Models [53.92832054643197]
We propose YAYI 2, including both base and chat models, with 30 billion parameters.
YAYI 2 is pre-trained from scratch on a multilingual corpus which contains 2.65 trillion tokens filtered by our pre-training data processing pipeline.
The base model is aligned with human values through supervised fine-tuning with millions of instructions and reinforcement learning from human feedback.
arXiv Detail & Related papers (2023-12-22T17:34:47Z) - Baichuan 2: Open Large-scale Language Models [51.56361715162972]
We present Baichuan 2, a series of large-scale multilingual language models containing 7 billion and 13 billion parameters, trained from scratch, on 2.6 trillion tokens.
Baichuan 2 matches or outperforms other open-source models of similar size on public benchmarks like MMLU, CMMLU, GSM8K, and HumanEval.
arXiv Detail & Related papers (2023-09-19T04:13:22Z) - How Far Can Camels Go? Exploring the State of Instruction Tuning on Open
Resources [117.6496550359768]
This work explores recent advances in instruction-tuning language models on a range of open instruction-following datasets.
We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets.
We evaluate them on their factual knowledge, reasoning, multilinguality, coding, and open-ended instruction following abilities.
arXiv Detail & Related papers (2023-06-07T19:59:23Z) - ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training
for Language Understanding and Generation [50.036392756981016]
GPT-3 has shown that scaling up pre-trained language models can further exploit their enormous potential.
A unified framework named ERNIE 3.0 was recently proposed for pre-training large-scale knowledge enhanced models.
ERNIE 3.0 outperformed the state-of-the-art models on various NLP tasks.
arXiv Detail & Related papers (2021-12-23T17:35:48Z) - DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with
Gradient-Disentangled Embedding Sharing [117.41016786835452]
This paper presents a new pre-trained language model, DeBERTaV3, which improves the original DeBERTa model.
vanilla embedding sharing in ELECTRA hurts training efficiency and model performance.
We propose a new gradient-disentangled embedding sharing method that avoids the tug-of-war dynamics.
arXiv Detail & Related papers (2021-11-18T06:48:00Z) - ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language
Understanding and Generation [25.430130072811075]
We propose a unified framework named ERNIE 3.0 for pre-training large-scale knowledge enhanced models.
It fuses auto-regressive network and auto-encoding network, so that the trained model can be easily tailored for both natural language understanding and generation tasks.
We trained the model with 10 billion parameters on a 4TB corpus consisting of plain texts and a large-scale knowledge graph.
arXiv Detail & Related papers (2021-07-05T16:54:59Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.