Model Details

Domain:

Language

Task:

Code generation

Model Access:

Open weights (unrestricted)

Citations:

766

AI Tools Usage

This model is commonly used behind the scenes in AI tools.

Introduction

Large language models (LMs) of code have recently shown tremendous promise in completing code and synthesizing code from natural language descriptions. However, the current state-of-the-art code LMs (e.g., Codex (Chen et al., 2021)) are not publicly available, leaving many questions about their model and data design decisions. We aim to fill in some of these blanks through a systematic evaluation of the largest existing models: Codex, GPT-J, GPT-Neo, GPT-NeoX-20B, and CodeParrot, across various programming languages. Although Codex itself is not open-source, we find that existing open-source models do achieve close results in some programming languages, although targeted mainly for natural language modeling. We further identify an important missing piece in the form of a large open-source model trained exclusively on a multi-lingual corpus of code. We release a new model, PolyCoder, with 2.7B parameters based on the GPT-2 architecture, which was trained on 249GB of code across 12 programming languages on a single machine. In the C programming language, PolyCoder outperforms all models including Codex. Our trained models are open-source and publicly available at this https URL, which enables future research and application in this area.

Benchmarking

FLOPs1.1e+21

Notes: "We use GPT-NeoX toolkit 11 to train the model efficiently in parallel with 8 Nvidia RTX 8000 GPUs on a single machine. The wall time used to train the largest 2.7B model is about 6 weeks" 8 * 130 TFLOP/s * 6 * 7 * 24 * 3600 * 0.3 (utilization) ~= 1.1e21

Training

Training Code AccessibilityMIT license for model weights https://huggingface.co/NinedayWang/PolyCoder-2.7B It seems that there is no pretraining code here: https://github.com/VHellendoorn/Code-LMs

HardwareNVIDIA Quadro RTX 8000

Size Notes: 249GB They trained on 39B tokens per Table 3, but I'm not sure how many epochs that is. May be <1.

Parameters

Parameters2700000000

Notes: 2.7B for largest model

Authors

Frank F. Xu, Uri Alon, Graham Neubig, Vincent J. Hellendoorn

Model Details

Domain:

Language

Task:

Code generation

Model Access:

Open weights (unrestricted)

Citations:

766

AI Tools Usage

This model is commonly used behind the scenes in AI tools.

Introduction

Benchmarking

FLOPs1.1e+21

Training

Training Code AccessibilityMIT license for model weights https://huggingface.co/NinedayWang/PolyCoder-2.7B It seems that there is no pretraining code here: https://github.com/VHellendoorn/Code-LMs

HardwareNVIDIA Quadro RTX 8000

Size Notes: 249GB They trained on 39B tokens per Table 3, but I'm not sure how many epochs that is. May be <1.

Parameters

Parameters2700000000

Notes: 2.7B for largest model

Authors

Frank F. Xu, Uri Alon, Graham Neubig, Vincent J. Hellendoorn

Carnegie Mellon University (CMU) | PolyCoder - Capabilities, Benchmarks and Use Cases

Top Tasks

Top Countries

Top Domains

Top Organizations

Top Categories

Top Collections

Platform

Top Tasks

Top Countries

Top Domains

Top Organizations

Top Categories

Top Collections

Platform

Model Details

AI Tools Usage

Introduction

Benchmarking

Training

Parameters

Authors

Top Tasks

Top Countries

Top Domains

Top Organizations

Top Categories

Top Collections

Platform

Model Details

AI Tools Usage

Introduction

Benchmarking

Training

Parameters

Authors