lucb1e 4 years ago

Since many people can't access the post anyhow, the short version of 'why' is: instagram keeps trying to block bibliogram (we can't have nice things), the author is tired of working around it, the author is also tired of fighting on a second front against the bots that try to scrape bibliogram. The bots also try hard to circumvent blocks that the author needed to put in place, rather than spending the time learning how to run their own instance and going to town on that.

Of course, it's open source so it's not like it's "discontinued" in the same sense as if gmail were to be announced to be discontinued. The official server will even remain online, just with current limitations (like viewing profiles being blocked by facebook) not being fixed for the foreseeable future.

  • dspillett 4 years ago

    > The bots also try hard to circumvent blocks that the author needed to put in place, rather than spending the time learning how to run their own instance and going to town on that.

    So a three-way arms race, with the author fighting a similar battle in parts against the bots that is being fought in the other direction with their bot and fb/insta.

    The enemy of my enemy may still be a pain in the arse for me!

    > Of course, it's open source so it's not like it's "discontinued"

    Aye, if there is enough value in it for someone else (who has the time and relevant skills) they'll pick up the reigns directly (if the author agrees) or in a fork – as happened when youtube-dl stagnated a bit.

    > The official server will even remain online

    And there are other instances. Or you can spin up your own if you don't mind (or want to try fix) the existing limitations & future breakages.

jeroenhd 4 years ago

Looks like there's a referrer filter for HN. Copy URL and open in a new tab, I guess?

My personal Bibliogram instance has been blocked for months. With Instagram blocked as part of my wider PiHole block on Facebook's domain and public Bibliogram instances shutting down soon, I guess I'm going to just ignore Instagram links from now on.

  • cercatrova 4 years ago

    I use an extension to strip referrer stuff from the URL, works well

    • lolinder 4 years ago

      Did it work here? The URL is clean, it's the request headers that give away the referer.

      • qjx 4 years ago

        i disabled header referrer in firefox's `about:configs` and it works for me

  • donkarma 4 years ago

    i just disable referer in my fierfox

    • gzer0 4 years ago

      And this just makes your browser fingerprint more unique

      • 3np 4 years ago

        One of us!

      • qjx 4 years ago

        why? i also disabled it thinking it'll make it better

        wouldn't it be the same as if someone just copy the link and open it in a new tab/window?

  • justsomehnguy 4 years ago

    I'm happily ignoring Twitter, Facebook, Instagram links all the time... until I need that post.

    Thankfully there is still no (and I hope there wouldn't be) tech related Instagram posts.

agluszak 4 years ago

I'm glad that Bibliogram exists, but it really shouldn't have to exist. Instagram/Meta should be legally required to provide an open API.

  • kQq9oHeAz6wLLS 4 years ago

    > Instagram/Meta should be legally required to provide an open API.

    Why?

    • atoav 4 years ago

      One argument would be: Because beyond a certain size your service becomes a public platfrom and you are no longer competing with other platforms through your features, but through the fact that you have an existing userbase. So there is A) public interest in accessing, archiving, displaying, etc. the things users do on your platform and B) forcing your service to become more open the bigger it becomes keeps the competition of services going.

      Edit: also see the relevant discourse in this other thread (especially that comment here): https://news.ycombinator.com/item?id=32686866

      • noitpmeder 4 years ago

        So have other services, i.e. PornHub and the like, crossed your threshold of public service?

        I've always viewed this argument as a very poor one. This essentially mandates that all very successful online enterprises become government arms. When does Instragram/Twitter/Facebook, once made a "public platform", become free of these new responsibilities? Is Twitter free to shut down tomorrow once they have public fiduciary responsibilities?

        I'm all in favor of additional regulation to encourage/reward/enforce "openness", especially with respect to data originating from users (photos, ...). But saying these companies have any responsibility to the social good is pure foolishness.

        • rhn_mk1 4 years ago

          > This essentially mandates that all very successful online enterprises become government arms.

          It doesn't. You can always split up before you grow too large, and remain private.

          There's precedent: once you're a monopoly, you have extra responsibilities to the social good, like splitting up or getting regulated.

          • ssizn 4 years ago

            Instagram has a monopoly on what exactly

            • aerzen 4 years ago

              It has dominant market share of photo-centric social platforms.

              In some smaller communities, it is also has a monopoly on social media in general and instant messaging.

              This is even more problematic, because such platforms have a positive feedback loop of user acquisition. Many users have no option but to use them, because all their contacts are using them.

              • ekianjo 4 years ago

                The definition of a tiny specific monopoly within a niche is meaningless. If you had 30 different factors to describe exactly what each platform does then each of them would be a monopoly in themselves.

                In any case Instagram is certainly not in a monopolistic position at the moment unless you try to twist the sense of words to fit a narrative.

            • rhn_mk1 4 years ago

              Instagram has a monopoly said who exactly?

        • atoav 4 years ago

          If you read my comment again you will find that I did not communicate this as my argument on purpose. So please don't frame it as such. Discussing something without necessarily having to share that opinion is hopefully still possible.

          Also: we are talking about Instagram, not about Pornhub — I guess the nature of the service would certainly play a role in that argument (e.g. think about messengers).

          What I do think myself is that beyond a certain size platforms should have it harder, should be held to higher standards, should fulfill more requirements etc. Because naturally bigger platforms will have it easier than their smaller competitors and it is in our interest to balance this out. Whether APIs are part of this, I don't know. Haven't given it much thought.

        • ad404b8a372f2b9 4 years ago

          All companies have a responsibility to the social good, it's not foolishness it's been the way things work since forever. Companies are responsible for their employees' well being, for their impact on the environment, and they have many responsibilities to the government, this is all enshrined in many many laws. See also all the companies that came together to produce Covid equipment at cost at the beginning of the pandemic.

    • grumbel 4 years ago

      It's just taking "Right to data portability"[1] a step further. We already have the ability to get data out of a service, but it's generally a slow and cumbersome process (e.g. giant .zip file that takes a day to create). So the next logical step would to require this process to be as seamless as possible, meaning real time import/export of selected bits of information, which could be accomplished by providing an API.

      Basically, if we ever want to get the Internet back into the users hand, we need to decouple the "service" and the "storage" parts. Companies shouldn't be allowed to hold user data hostage.

      [1] https://gdpr-info.eu/art-20-gdpr/

    • baq 4 years ago

      Why not?

  • Bukhmanizer 4 years ago

    I don’t mind requiring open APIs but can we all agree that this is pretty much in direct opposition to a lot of the privacy concerns that people here talk about?

    The more able I am to interact with an interface programmatically, the more easily I can exploit it and turn it into a giant database.

    • Nextgrid 4 years ago

      The privacy problem with Instagram isn't what people post on it (they want it to be seen publicly), it's what the disgusting company behind it collects in the background as you interact with it (who may be a non-Instagram user and just need to visit an Instagram link you've been sent).

      People don't use Bibliogram (or other alternatives) to access unauthorized data (all it displays is the same data visible on the web or via the app), it's to be able to browse without being stalked by Facebook and avoid their dark patterns.

      • Bukhmanizer 4 years ago

        People might not want to be seen publicly by everyone, which is why they allow users to choose who to show their posts to. But it’s not all that hard to exploit this by creating a popular page that lots of people end up sharing some data with. With permissive enough APIs, you instantly have a dataset.

        This is more or less how the Cambridge Analytica fiasco started. And yes, this was seen as a privacy problem.

        • 411111111111111 4 years ago

          Do keep in mind that the people that are calling for a default API here aren't the same people that would've complained about the privacy implications of said API.

          This is a diverse Forum and people have various opinions.

          • Bukhmanizer 4 years ago

            > Do keep in mind that the people that are calling for a default API here aren't the same people that would've complained about the privacy implications of said API.

            To be honest I think they would have. I just don’t think they’ve fully thought through the consequences. Worse, maybe they have thought about the consequences, but are ok with being outraged at both ends of things. Lots of people get mad and try and skirt any security/password protection measure, but also get just as mad when they inevitably get hacked because they’ve skirted all the security measures. But even after that still won’t use the security measures because of the “inconvenience”.

        • Nextgrid 4 years ago

          > People might not want to be seen publicly by everyone

          Private accounts (where you have to approve followers manually) exist and as far as I know nobody is calling for making those public. A hypothetical API would still require the credentials of an approved account before returning data of such accounts.

          > But it’s not all that hard to exploit this by creating a popular page that lots of people end up sharing some data with

          I don't understand what this has to do with an API? You can create a phishing page now.

          > This is more or less how the Cambridge Analytica fiasco started. And yes, this was seen as a privacy problem.

          Cambridge Analytica created a malicious website that requested access to people's accounts and they granted it. Not only is it on idiots that approved the OAuth consent form, but it doesn't seem like it's got anything to do with what the parent comment was proposing? As far as I know all he was calling for is to provide an API that returns the public data that you can already get via the official web UI or app.

      • ssizn 4 years ago

        If everybody accesses Instagram from bibliogram, bypassing ads and data collection techniques, who’s going to pay for the service? I know that your answer will be “so close it”, but that’s hardly what the users of the platform want.

        • Nextgrid 4 years ago

          You can target ads based on public data such as posts on your profile and data you voluntarily submit (such as the accounts you follow, etc).

    • AlecSchueler 4 years ago

      You can still have the same permissions on your API.

  • jacooper 4 years ago

    If the digital markets act expands to social media, it will happen.

  • annadane 4 years ago

    Or legally required to provide a chronological feed, which even logged-in users don't get

    • tech234a 4 years ago

      Have you tried clicking the Instagram logo at the top left lately? It should open a drop-down menu where you can pick “following”, though I’m not sure if this has rolled out to everyone yet.

prvit 4 years ago

lol, OP completely missed out on the biggest IG ratelimit bypass. You could just use IPv6 and make literally millions of requests per second from a single server with a /48. (Oh! And before that you could just spoof X-Forwarded-For)

I printed money using this to automatically take good usernames as they became available.

  • carabiner 4 years ago

    How much did you make total?

    • prvit 4 years ago

      A little over $2M, but that number keeps growing because I still have loads of usernames despite instagram patching the ratelimit bypass.

      • ozarker 4 years ago

        Jesus, more power to you I guess

      • carabiner 4 years ago

        What are you doing now? Do you still have a day job?

        • prvit 4 years ago

          I've never had a day job, I've spent almost two decades funding my lifestyle with silly projects like this. Everything from large scale world of warcraft botting to instagram registrations.

          • carabiner 4 years ago

            Wow. Where do you live? What's your net worth, if you're willing to reveal?

          • anonymouse008 4 years ago

            [edit] I’ll heed

            • encryptluks2 4 years ago

              They are a scammer. Don't be surprised if you get scammed.

              • prvit 4 years ago

                Unnecessarily harsh, but I'm not interested in selling trainings anyway.

          • carabiner 4 years ago

            Man, do an AMA. This is more fascinating than 90% of what's posted here.

            If you're able to find an edge and exploit it in numerous disparate markets, I bet you'd be amazing in algotrading.

            • from 4 years ago

              It may interest you to know that the guy you are replying to is Julius “zeekill” Kivimaki of Lizard Squad fame.

          • FearlessNebula 4 years ago

            Are you not worried about running into legal trouble with these things? It seems like a gray area.

            Also, what platform do you use to sell these things? How do you know the buyer won’t just take the credentials and not send payment? Or how do they know you won’t take payment and refuse to hand over the credentials?

            • prvit 4 years ago

              > Are you not worried about running into legal trouble with these things? It seems like a gray area.

              No, not really. I get the occasional nastygram from Hogan Lovells, but just shrug them off. I don’t do anything interesting enough for them to bother litigating.

              > Also, what platform do you use to sell these things?

              There are many, swapd.co is a good one.

              > Or how do they know you won’t take payment and refuse to hand over the credentials?

              Reputation, there are also a plenty of reputable escrow services available.

              • FearlessNebula 4 years ago

                Very cool. I’ll echo what others have said, it would be very cool to see some sort of how to or read an AMA, about the strategies that maybe no longer work (since obviously you wouldn’t want to reveal something that is still making you money). I’d definitely read a blog post about how you used to make WoW bots or automate creating accounts

      • shard 4 years ago

        I am guessing that the $2M is from selling the usernames? Are the buyers individuals who want a good username, or bot farms who just need a large number of accounts?

        • prvit 4 years ago

          Buyers are individuals who want good usernames. People will pay like $5k for any random 2 character username. I've sold words for 50.

          • shard 4 years ago

            Interesting. Where does one go to buy usernames for social networks?

  • cosmojg 4 years ago

    You brilliant bastard.

  • nubela 4 years ago

    Is there an email I can reach out to you at? Not for this particular use-case of IG but would love to pick you brain on rate limit bypasses.

  • sprite 4 years ago

    How do you get a /48? Tunnelbroker? Can you set external IP per request? Curious how this works.

    • prvit 4 years ago

      You could just get an allocation from RIPE for example, they'll give you a /29 without any justification required. Pretty much all dedicated server hosts will have a plenty of IPv6 space they can give you for a few $ per month.

      Tunnels would've been too slow for my purposes.

      > Can you set external IP per request?

      Yeah, of course. You can just use the IP_FREEBIND socket option.

  • icehawk 4 years ago

    Knowing someone who ran a public bibliogram instance, using IPv6 and rotating through addresses of a cellular provider's /32 it was, at the barest minimum, possible, but really annoying as Facebook would lock down any /64 once it felt it had looked at too many profile pages, and the limits were just as tight, if not more so, as the v4 limits.

    • prvit 4 years ago

      Yes, but you could easily defeat this limit by switching to a /48, which gives you 65536 /64 subnets. Of course, even bigger IPv6 allocations are easy to get.

  • ddevault 4 years ago

    Man, fuck that. Name scalping is wrong. This shit is part of why platforms lock down in the first place.

TobyTheDog123 4 years ago

Anti-scraping technology is absolutely ridiculous. Modern example: TikTok. Have you seen webmsssdk.js?

They simply don't want users to be able to access their own data and use it elsewhere in the internet that they can't see, control, and sell.

This is exaggerated by services that login to your accounts to scrape your data for you being illegal under the CFAA.

What a shame for the web.

  • jbk 4 years ago

    > Anti-scraping technology is absolutely ridiculous. Modern example: TikTok. Have you seen webmsssdk.js?

    What is webmssdk doing?

    • bashonly 4 years ago

      for all of the requests to their api, it constructs a very long query string that is a basic fingerprint of your browser/os. then the query needs to be signed (or else the api returns a captcha), which is done by a blob of encrypted/obfuscated-beyond-recovery JS that uses a more comprehensive browser fingerprint to validate the query string and generate a token that is appended to the query. there's another required query param which involves an xhr to phone home so that presumably the fingerprints can be checked server-side. finally all of your fingerprint data is sent to their api where it lives happily ever after and you get some 15 second videos to watch

  • ssizn 4 years ago

    >They simply don't want users to be able to access their own data and use it elsewhere in the internet that they can't see, control, and sell.

    That’s false, you can download a copy of your own data from settings.

3np 4 years ago

Reminder to anyone who might want to peruse or fork the source in the future: Sources still available at https://sr.ht/~cadence/bibliogram/sources but as with anything may become unavailable at any point.

fallingknife 4 years ago

Referrer filter, so I right clicked and "open link in new tab" but still referrer filter. Don't like that the browser sends that info with that option. That should be equivalent to copy/pasting the URL.

  • woojoo666 4 years ago

    Also why the HN filter in the first place? Seems like a vendetta against HN, especially given the message

    • bobsmooth 4 years ago

      Some people hate free publicity.

unicornporn 4 years ago

It's been nonfunctional for at least over a year. Changes to Instagram made impossible to bypass the wall protecting the garden (or landfill).

technonerd 4 years ago

I ran into the private profile problem with bibliogram ALL THE TIME. Even tho I was friends with them on instagram. So I just stop using it. It was broken most of the time anyways. It was a fun 3rd party unoffical API frontend that I added to my list of other ones I could self host.

That amount of hostility to scrapping and 3rd party clients feels like that they want to be in control of the data not you.

lapinot 4 years ago

Thank you cadence for all the hard work and inspiration. This wasn't for nothing. I would bet on a newleaf-style reboot (with gallery-dl, to tackle onto a bigger community) if i had to. I might try my hand at it at some point but i'm afraid i'm less dedicated than you are.

rektide 4 years ago

I found a really awesome "public instance" vps node that had a bunch of utils & anonymous public interfaces webapps (nitter,...) running on it for all.

They also had a list of things they dont support, with reasons why. Meta interfacing were listed (including to bibliogram), and typically with links to fire-bot takedown requests to kingdom come. No one else has anything like this. Very very unique company.

NoboruWataya 4 years ago

I self-hosted this for a while but found that I got blocked/severely limited as soon as I tried to do basically anything. I'm sure there were things I could have done to get around the limitations that very quickly got imposed, but it wasn't worth it to me. If you don't like Instagram's UI/tracking/privacy policy/closedness then really the only sustainable solution is to stop using it, in any guise.

Bibliogram was a great project but I think Instagram is too aggressive about trying to protect its walled garden. It's a losing battle. Easier to just realise that, for most people, Instagram doesn't add that much to your life anyway.

nomilk 4 years ago

Never used an OSS frontend before. Any other open source frontends worth trying?

  • hadrien01 4 years ago

    Nitter is a frontend for Twitter. It's so much faster, and doesn't block you if you don't have a Twitter account.

    I don't use an alternative Web frontend for YouTube, but on Android I use NewPipe.

  • black_puppydog 4 years ago

    Libreddit!

    And like sibling said, nitter.

    Both are trivially easy to self host a private instance and with privacy redirect or a similar extension they just became my default way of opening twitter/reddit links.

    • qjx 4 years ago

      teddit.net is a good reddit frontend

      • black_puppydog 3 years ago

        yeah if you want to actually interact. but 99% of the time I just want to read, and libreddit is perfect for that

annadane 4 years ago

Take down some of the most egregious pages and disinformation? No, we (Facebook/Meta) will waste our time shutting down anything remotely challenging to our monstrosity we're trying to conquer the world with (see also: shutting down the NYU Ad Observatory, or any countless other examples)

  • Bakary 4 years ago

    What incentive have we given them to stop?

    • annadane 4 years ago

      Incentive? How about "don't be monopolistic assholes"?

      • austinjp 4 years ago

        That's more of a suggestion than an incentive, though.

      • Bakary 4 years ago

        I don't really see the point of appealing to morality or to what ought to happen. They're not going to relinquish power willingly. Few individual do, and for groups it's vanishingly rare.

        • lapinot 4 years ago

          "don't be monopolistic assholes" is in the constitution and laws of several countries, it should not be an appeal to morality. The real problem is that at any point anybody does a move constraining internet corps, they get mocked by everyone as "regulating things to death". Yes that's the goal. Their death, in their current form at least.

          • Bakary 4 years ago

            That might be true in the anglophone world, but plenty of other populations aren't hostile to the idea of regulating the internet, and I don't mean the totalitarian countries. The GDPR is a ray of hope in this regard.

            What prompted me to mention incentives in the first place is that the GP's framing of the problem is a common narcotic. Online is/ought-type complaining reduces the potential for action since the emotional need for signaling is already fulfilled. It's a trap I've fallen into many times myself. In itself, it's more dangerous at scale than any public mockery of regulation.

          • pessimizer 4 years ago

            > "don't be monopolistic assholes" is in the constitution and laws of several countries,

            No it's not. People who formulate laws are specific, unless they're either creating a law that is intended to be ignored, or if they're creating a law meant for selective enforcement.

            You can check every constitution and law code all day, you're never going to find anything telling anyone not to impede the attempts of an application not to constantly scrape their webpage. This kind of talk is purposeless.

CyborgCabbage 4 years ago

We still have teddit.net and nitter.net

dspillett 4 years ago

> Before we start: If Bibliogram has been helpful to you, please consider making a donation!

Before asking me to consider a donation, consider cutting the silly referrer based moan about the the site that referred me to yours so I have to jump through a hoop to see the content.

  • textadventure 4 years ago

    It stands to reason that they don't care for HN referred donations as the hoop seems to be very deliberately put in place for a reason.

    • blitzar 4 years ago

      Rather than a pissy referrer filter, why not redirect to a page full of ads.

      • kixiQu 3 years ago

        Which is pissier: the filter, or the complaints about it?

        • blitzar 3 years ago

          100% my complaint

    • dspillett 4 years ago

      A reason that is not explained, so I'm going to assume is something petty. There are a couple of instances of sites doing this (for HN and other referral sources) because people said things they didn't like about them (or just in general) on the platform.