Topic: reinforcement-learning-from-human-feedback Goto Github
Some thing interesting about reinforcement-learning-from-human-feedback
Some thing interesting about reinforcement-learning-from-human-feedback
reinforcement-learning-from-human-feedback,LMRax is a framework built on JAX to train transformers language models by reinforcement learning, along with the reward model training.
Organization: almost-intelligence
reinforcement-learning-from-human-feedback,annotated tutorial of the huggingface TRL repo for reinforcement learning from human feedback connecting equations from PPO and GAE to the lines of code in the pytorch implementation
User: clam004
reinforcement-learning-from-human-feedback,[TSMC] Ask-AC: An Initiative Advisor-in-the-Loop Actor-Critic Framework
User: liushunyu
Home Page: https://ieeexplore.ieee.org/abstract/document/10210582
reinforcement-learning-from-human-feedback,Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback
User: nlp-uoregon
reinforcement-learning-from-human-feedback,An Easy-to-use, Scalable and High-performance RLHF Framework (Support 70B+ full tuning & LoRA & Mixtral & KTO)
Organization: openllmai
Home Page: https://huggingface.co/OpenLLMAI
reinforcement-learning-from-human-feedback,Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Organization: pku-alignment
Home Page: https://pku-beaver.github.io
reinforcement-learning-from-human-feedback,A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
Organization: tatsu-lab
Home Page: https://arxiv.org/abs/2305.14387
reinforcement-learning-from-human-feedback,A repo for RLHF training and BoN over LLMs, with support for reward model ensembles.
User: tlc4418
Home Page: https://arxiv.org/abs/2310.02743
reinforcement-learning-from-human-feedback,Shaping Language Models with Cognitive Insights
Organization: xplainmind
reinforcement-learning-from-human-feedback,RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback
User: ymetz
Home Page: https://rlhfblender.readthedocs.io/en/latest/
reinforcement-learning-from-human-feedback,Summaries of papers related to the alignment problem in NLP
User: ymnseol
A declarative, efficient, and flexible JavaScript library for building user interfaces.
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
An Open Source Machine Learning Framework for Everyone
The Web framework for perfectionists with deadlines.
A PHP framework for web artisans
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
Some thing interesting about web. New door for the world.
A server is a program made to process requests and deliver data to clients.
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
Some thing interesting about visualization, use data art
Some thing interesting about game, make everyone happy.
We are working to build community through open source technology. NB: members must have two-factor auth.
Open source projects and samples from Microsoft.
Google ❤️ Open Source for everyone.
Alibaba Open Source for everyone
Data-Driven Documents codes.
China tencent open source team.