Maybe another Jev based on BERT
TypeSafe hasn’t published Jev’s architecture, its weights, or a technical paper. They’ve stated that the model is based on a transformer, that it’s trained only on synthetic data, and that it isn’t autoregressive, meaning it doesn’t produce the answer one token at a time. According to TypeSafe, Jev uses a new architecture with a parallel sampler, produces all the answers in a single query, and evaluates the questions asked on the same state in parallel and independently. The probabilities are calibrated with a post-training step that TypeSafe calls RLCD (Reinforcement Learning for Calibrated Decisions), which compares them with the observed outcomes and corrects them. ...