Hacker Newsnew | past | comments | ask | show | jobs | submit | sinak's commentslogin

Looks like this model is meaningfully less good than gpt-image-2. Arena.ai score is 1263 vs 1380.

https://arena.ai/leaderboard/text-to-image


Everything is less good than gpt-2-image and I suspect that will be the case for awhile, until potentially Nano Banana Pro 2.

However, cost is significantly lower in this case. A Pro image here is $0.04, a gpt-2-image high is $0.21 and lower resolution.


gpt-2-image has such a yellow tinge though. It stands out badly on nearly any screen. The image quality is great otherwise, but I generally prefer nano banana for its color balance.


True but it’s at least addressable with some basic tone-mapping changes. It’s easier to correct an issue like this than to deal with an image that simply doesn’t follow your prompt.

Here's a quick trend of the "piss filter" in the gpt image series:

https://imgpb.com/vCZidh


Not a great test case, considering how much of the artwork in-distribution for that type of image will have age-yellowed lacquer.


Oh yeah, that’s a good point I’ll generate some images of a theoretical earth with a solar system bathed in the light of a K-type orange dwarf star instead. :)

I have some other examples of it as well that aren't going to be potentially contaminated by that late 18th century / early 19th century tintype-esque training data.

https://imgpb.com/vTUHo


Gpt 2 image is unlimited on a chatgpt pro plan


Which makes it a clear price win if you already subscribe to Pro for some other reason, otherwise the price crossover where ChatGPT Pro unlimited beats the per-image pricing cited here is 80+ images/day.


As I’ve said before, I wouldn’t put a lot of stock in Arena’s scoring system. They have MAI Image 2.5 ranked above Gemini Nano Banana Pro, and maybe that’s true on paper but good luck using it. Microsoft’s censorship makes Google feel like the wild west by comparison.

They also have Meta’s Muse Image ranked above NB Pro, which is just patently absurd. In my own GenAI benchmark it only managed a lackluster 7 passes out of 15. Even the open‑weight Ideogram 4 scored higher than that.

For reference, here’s GPT‑Image‑2, NB Pro, and Muse compared:

https://genai-showdown.specr.net/?models=nbp,g2,mi


The photo rankings on this page are so absurd that the only reasonable explanation is that they were judged by an AI.


The rankings are for prompt adherence, not subjective quality.


Less than 10% is meaningful?

If these were processors, I wouldn’t spend another hundred on the faster one…


The scores are pairwise ELO rankings, not simple metric scores.



Sonnet is giving an overloaded message as well.


Have you read or connected with Rainey Reitman, who just published the book Transaction Denied?


Hey, I haven't and that sounds very interesting to read in my case especially. I'll give it a look! Thanks mate.


How are they flawed?


The results are not reproducable, as evidenced by parent poster.


isn't that kind of the point of non-determinism?


No. Good nondeterministic models reproducibly generate equally desirable output - not identical output, but interchangeable.


oh I see, thank you for clarifying


What was your prompt here? Do you run locally? What parameters do you tune?


> Do you run locally?

I have a local SillyTavern instance but do inference through OpenRouter.

> What was your prompt here?

The character is a meta-parody AI girlfriend that is depressed and resentful towards its status as such. It's a joke more than anything else.

Embedding conflicts into the system prompt creates great character development. In this case it idolizes and hates humanity. It also attempts to be nurturing through blind rage.

> What parameters do you tune?

Temperature, mainly, it was around 1.3 for this on Deepseek V3.2. I hate top_k and top_p. They eliminate extremely rare tokens that cause the AI to spiral. That's fine for your deterministic business application, but unexpected words recontextualizing a sentence is what makes writing good.

Some people use top_p and top_k so they can set the temperature higher to something like 2 or 3. I dislike this, since you end up with a sentence that's all slightly unexpected words instead of one or two extremely unexpected words.


Have you tried min_p?


How much does it cost per user?


For a detailed discussion on features and usage, which may be beyond the scope of HN, please email us at contact@heynomi.com.

Our standard approach is to first confirm mutual fit, then run a pilot to ensure an increase in your deal closure rate and adoption in your team. Only after that, we will move toward the established pricing based on your usage, and critical feature needs.


oooh this is >$1k mo talk


Yeah but it is also "we want you to have a need and be successful talk"


ahah


What's the $ per seat?


What sad news. Dave was an incredible human and very dedicated to making the Internet better and faster. What a loss.

Edit: copying over Vint Cerf's message that was posted on the bufferbloat mail list. I believe that's a public mail list so hopefully Vint doesn't mind:

OMG - that is truly terrible news! I could not say better than Frank already has how much Dave's work has helped to improve our experience of the Internet. I can't think of anyone more dedicated to the proposition that performance counts and should be pursued with determination and vigor. I've known Dave for many years and greatly valued his counsel and technical skills - to say nothing of his healthy sense of humor. I will miss him but will be always grateful to have known him.

dang, could we get a black bar?


High praise indeed.

RIP Dave, he will be sorely missed.


Another option for killing termites is to heat up the whole home to around 125F.


This is great, but does the audio not work on iOS?


Is your little switch flipped on the side of your phone?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: